SearcharxivSearch

arXiv subjects

Kevin Schawinski

Publications and source records attributed to Kevin Schawinski.

At least 19 recordsLinked to original sources

Sparks of Science: Hypothesis Generation Using Structured Paper Data

Generating novel and creative scientific hypotheses is a cornerstone in achieving Artificial General Intelligence. Large language and reasoning models have the potential to aid in the systematic creation, selection, and validation of scientifically informed hypotheses. However, current foundation models often struggle to produce scientific ideas that are both novel and feasible. One reason is the lack of a dedicated dataset that frames Scientific Hypothesis Generation (SHG) as a Natural Language Generation (NLG) task. In this paper, we introduce HypoGen, the first dataset of approximately 5500 structured problem-hypothesis pairs extracted from top-tier computer science conferences structured with a Bit-Flip-Spark schema, where the Bit is the conventional assumption, the Spark is the key insight or conceptual leap, and the Flip is the resulting counterproposal. HypoGen uniquely integrates an explicit Chain-of-Reasoning component that reflects the intellectual process from Bit to Flip. We demonstrate that framing hypothesis generation as conditional language modelling, with the model fine-tuned on Bit-Flip-Spark and the Chain-of-Reasoning (and where, at inference, we only provide the Bit), leads to improvements in the overall quality of the hypotheses. Our evaluation employs automated metrics and LLM judge rankings for overall quality assessment. We show that by fine-tuning on our HypoGen dataset we improve the novelty, feasibility, and overall quality of the generated hypotheses. The HypoGen dataset is publicly available at huggingface.co/datasets/UniverseTBD/hypogen-dr1.

cs.CL

A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models

Hypothesis generation is a fundamental step in scientific discovery, yet it is increasingly challenged by information overload and disciplinary fragmentation. Recent advances in Large Language Models (LLMs) have sparked growing interest in their potential to enhance and automate this process. This paper presents a comprehensive survey of hypothesis generation with LLMs by (i) reviewing existing methods, from simple prompting techniques to more complex frameworks, and proposing a taxonomy that categorizes these approaches; (ii) analyzing techniques for improving hypothesis quality, such as novelty boosting and structured reasoning; (iii) providing an overview of evaluation strategies; and (iv) discussing key challenges and future directions, including multimodal integration and human-AI collaboration. Our survey aims to serve as a reference for researchers exploring LLMs for hypothesis generation.

cs.CL

Automatic Machine Learning Framework to Study Morphological Parameters of AGN Host Galaxies within $z < 1.4$ in the Hyper Supreme-Cam Wide Survey

We present a composite machine learning framework to estimate posterior probability distributions of bulge-to-total light ratio, half-light radius, and flux for Active Galactic Nucleus (AGN) host galaxies within $z<1.4$ and $m<23$ in the Hyper Supreme-Cam Wide survey. We divide the data into five redshift bins: low ($0<z<0.25$), mid ($0.25<z<0.5$), high ($0.5<z<0.9$), extra ($0.9<z<1.1$) and extreme ($1.1<z<1.4$), and train our models independently in each bin. We use PSFGAN to decompose the AGN point source light from its host galaxy, and invoke the Galaxy Morphology Posterior Estimation Network (GaMPEN) to estimate morphological parameters of the recovered host galaxy. We first trained our models on simulated data, and then fine-tuned our algorithm via transfer learning using labeled real data. To create training labels for transfer learning, we used GALFIT to fit $\sim 20,000$ real HSC galaxies in each redshift bin. We comprehensively examined that the predicted values from our final models agree well with the GALFIT values for the vast majority of cases. Our PSFGAN + GaMPEN framework runs at least three orders of magnitude faster than traditional light-profile fitting methods, and can be easily retrained for other morphological parameters or on other datasets with diverse ranges of resolutions, seeing conditions, and signal-to-noise ratios, making it an ideal tool for analyzing AGN host galaxies from large surveys coming soon from the Rubin-LSST, Euclid, and Roman telescopes.

astro-ph.GA

pathfinder: A Semantic Framework for Literature Review and Knowledge Discovery in Astronomy

The exponential growth of astronomical literature poses significant challenges for researchers navigating and synthesizing general insights or even domain-specific knowledge. We present Pathfinder, a machine learning framework designed to enable literature review and knowledge discovery in astronomy, focusing on semantic searching with natural language instead of syntactic searches with keywords. Utilizing state-of-the-art large language models (LLMs) and a corpus of 350,000 peer-reviewed papers from the Astrophysics Data System (ADS), Pathfinder offers an innovative approach to scientific inquiry and literature exploration. Our framework couples advanced retrieval techniques with LLM-based synthesis to search astronomical literature by semantic context as a complement to currently existing methods that use keywords or citation graphs. It addresses complexities of jargon, named entities, and temporal aspects through time-based and citation-based weighting schemes. We demonstrate the tool's versatility through case studies, showcasing its application in various research scenarios. The system's performance is evaluated using custom benchmarks, including single-paper and multi-paper tasks. Beyond literature review, Pathfinder offers unique capabilities for reformatting answers in ways that are accessible to various audiences (e.g. in a different language or as simplified text), visualizing research landscapes, and tracking the impact of observatories and methodologies. This tool represents a significant advancement in applying AI to astronomical research, aiding researchers at all career stages in navigating modern astronomy literature.

astro-ph.IM

AstroLLaMA-Chat: Scaling AstroLLaMA with Conversational and Diverse Datasets

We explore the potential of enhancing LLM performance in astronomy-focused question-answering through targeted, continual pre-training. By employing a compact 7B-parameter LLaMA-2 model and focusing exclusively on a curated set of astronomy corpora -- comprising abstracts, introductions, and conclusions -- we achieve notable improvements in specialized topic comprehension. While general LLMs like GPT-4 excel in broader question-answering scenarios due to superior reasoning capabilities, our findings suggest that continual pre-training with limited resources can still enhance model performance on specialized topics. Additionally, we present an extension of AstroLLaMA: the fine-tuning of the 7B LLaMA model on a domain-specific conversational dataset, culminating in the release of the chat-enabled AstroLLaMA for community use. Comprehensive quantitative benchmarking is currently in progress and will be detailed in an upcoming full paper. The model, AstroLLaMA-Chat, is now available at https://huggingface.co/universeTBD, providing the first open-source conversational AI tool tailored for the astronomy community.

astro-ph.IM

AstroLLaMA: Towards Specialized Foundation Models in Astronomy

Large language models excel in many human-language tasks but often falter in highly specialized domains like scholarly astronomy. To bridge this gap, we introduce AstroLLaMA, a 7-billion-parameter model fine-tuned from LLaMA-2 using over 300,000 astronomy abstracts from arXiv. Optimized for traditional causal language modeling, AstroLLaMA achieves a 30% lower perplexity than Llama-2, showing marked domain adaptation. Our model generates more insightful and scientifically relevant text completions and embedding extraction than state-of-the-arts foundation models despite having significantly fewer parameters. AstroLLaMA serves as a robust, domain-specific model with broad fine-tuning potential. Its public release aims to spur astronomy-focused research, including automatic paper summarization and conversational agent development.

astro-ph.IM

Using Machine Learning to Determine Morphologies of $z<1$ AGN Host Galaxies in the Hyper Suprime-Cam Wide Survey

We present a machine-learning framework to accurately characterize morphologies of Active Galactic Nucleus (AGN) host galaxies within $z<1$. We first use PSFGAN to decouple host galaxy light from the central point source, then we invoke the Galaxy Morphology Network (GaMorNet) to estimate whether the host galaxy is disk-dominated, bulge-dominated, or indeterminate. Using optical images from five bands of the HSC Wide Survey, we build models independently in three redshift bins: low $(0<z<0.25)$, medium $(0.25<z<0.5)$, and high $(0.5<z<1.0)$. By first training on a large number of simulated galaxies, then fine-tuning using far fewer classified real galaxies, our framework predicts the actual morphology for $\sim$ $60\%-70\%$ host galaxies from test sets, with a classification precision of $\sim$ $80\%-95\%$, depending on redshift bin. Specifically, our models achieve disk precision of $96\%/82\%/79\%$ and bulge precision of $90\%/90\%/80\%$ (for the 3 redshift bins), at thresholds corresponding to indeterminate fractions of $30\%/43\%/42\%$. The classification precision of our models has a noticeable dependency on host galaxy radius and magnitude. No strong dependency is observed on contrast ratio. Comparing classifications of real AGNs, our models agree well with traditional 2D fitting with GALFIT. The PSFGAN+GaMorNet framework does not depend on the choice of fitting functions or galaxy-related input parameters, runs orders of magnitude faster than GALFIT, and is easily generalizable via transfer learning, making it an ideal tool for studying AGN host galaxy morphology in forthcoming large imaging survey.

astro-ph.GA

BASS XXV: DR2 Broad-line Based Black Hole Mass Estimates and Biases from Obscuration

We present measurements of broad emission lines and virial estimates of supermassive black hole masses ($M_{BH}$) for a large sample of ultra-hard X-ray selected active galactic nuclei (AGNs) as part of the second data release of the BAT AGN Spectroscopic Survey (BASS/DR2). Our catalog includes $M_{BH}$ estimates for a total 689 AGNs, determined from the H$α$, H$β$, $MgII\lambda2798$, and/or $CIV\lambda1549$ broad emission lines. The core sample includes a total of 512 AGNs drawn from the 70-month Swift/BAT all-sky catalog. We also provide measurements for 177 additional AGNs that are drawn from deeper Swift/BAT survey data. We study the links between $M_{BH}$ estimates and line-of-sight obscuration measured from X-ray spectral analysis. We find that broad H$α$ emission lines in obscured AGNs ($\log (N_{\rm H}/{\rm cm}^{-2})> 22.0$) are on average a factor of $8.0_{-2.4}^{+4.1}$ weaker, relative to ultra-hard X-ray emission, and about $35_{-12}^{~+7}$\% narrower than in unobscured sources (i.e., $\log (N_{\rm H}/{\rm cm}^{-2}) < 21.5$). This indicates that the innermost part of the broad-line region is preferentially absorbed. Consequently, current single-epoch $M_{BH}$ prescriptions result in severely underestimated ($>$1 dex) masses for Type 1.9 sources (AGNs with broad H$α$ but no broad H$β$) and/or sources with $\log (N_{\rm H}/{\rm cm}^{-2}) > 22.0$. We provide simple multiplicative corrections for the observed luminosity and width of the broad H$α$ component ($L[{\rm b}{\rm H}α]$ and FWHM[bH$α$]) in such sources to account for this effect, and to (partially) remedy $M_{BH}$ estimates for Type 1.9 objects. As key ingredient of BASS/DR2, our work provides the community with the data needed to further study powerful AGNs in the low-redshift Universe.

astro-ph.GA

BASS XXVI: DR2 Host Galaxy Stellar Velocity Dispersions

We present new central stellar velocity dispersions for 484 Sy 1.9 and Sy 2 from the second data release of the Swift/BAT AGN Spectroscopic Survey (BASS DR2). This constitutes the largest study of velocity dispersion measurements in X-ray selected, obscured AGN with 956 independent measurements of the Ca H+K and Mg b region (3880-5550A) and the Ca triplet region (8350-8730A) from 642 spectra mainly from VLT/Xshooter or Palomar/DoubleSpec. Our sample spans velocity dispersions of 40-360 km/s, corresponding to 4-5 orders of magnitude in black holes mass (MBH=10^5.5-9.6 Msun), bolometric luminosity (LBol~10^{42-46 ergs/s), and Eddington ratio (L/Ledd~10^{-5}-2). For 281 AGN, our data provide the first published central velocity dispersions, including 6 AGN with low mass black holes (MBH=10^5.5-6.5 Msun), discovered thanks to our high spectral resolution observations (sigma~25 km/s). The survey represents a significant advance with a nearly complete census of hard-X-ray selected obscured AGN with measurements for 99% of nearby AGN (z<0.1) outside the Galactic plane. The BASS AGN have higher velocity dispersions than the more numerous optically selected narrow line AGN (i.e., ~150 vs. ~100 km/s), but are not biased towards the highest velocity dispersions of massive ellipticals (i.e., >250 km/s). Despite sufficient spectral resolution to resolve the velocity dispersions associated with the bulges of small black holes (~10^4-5 Msun), we do not find a significant population of super-Eddington AGN. Using estimates of the black hole sphere of influence, direct stellar and gas black hole mass measurements could be obtained with existing facilities for more than ~100 BASS AGN.

astro-ph.GA

BASS XXII: The BASS DR2 AGN Catalog and Data

We present the AGN catalog and optical spectroscopy for the second data release of the Swift BAT AGN Spectroscopic Survey (BASS DR2). With this DR2 release we provide 1425 optical spectra, of which 1181 are released for the first time, for the 858 hard X-ray selected AGN in the Swift BAT 70-month sample. The majority of the spectra (813/1425, 57%) are newly obtained from VLT/Xshooter or Palomar/Doublespec. Many of the spectra have both higher resolution (R>2500, N~450) and/or very wide wavelength coverage (3200-10000 A, N~600) that are important for a variety of AGN and host galaxy studies. We include newly revised AGN counterparts for the full sample and review important issues for population studies, with 44 AGN redshifts determined for the first time and 780 black hole mass and accretion rate estimates. This release is spectroscopically complete for all AGN (100%, 858/858) with 99.8% having redshift measurements (857/858) and 96% completion in black hole mass estimates of unbeamed AGN (outside the Galactic plane). This AGN sample represents a unique census of the brightest hard X-ray selected AGN in the sky, spanning many orders of magnitude in Eddington ratio (Ledd=10^-5-100), black hole mass (MBH=10^5-10^10 Msun), and AGN bolometric luminosity (Lbol=10^40-10^47 ergs/s).

astro-ph.GA

BAT AGN Spectroscopic Survey XXI: The Data Release 2 Overview

The BAT AGN Spectroscopic Survey (BASS) is designed to provide a highly complete census of the key physical parameters of supermassive black holes (SMBHs) that power local active galactic nuclei (AGN) (z<0.3), including their bolometric luminosity, black hole mass, accretion rates, and line-of-sight gas obscuration, and the distinctive properties of their host galaxies (e.g., star formation rates, masses, and gas fractions). We present an overview of the BASS data release 2 (DR2), an unprecedented spectroscopic survey in spectral range, resolution, and sensitivity, including 1449 optical (3200-10000 A) and 233 NIR (1-2.5 um) spectra for the brightest 858 ultra-hard X-ray (14-195 keV) selected AGN across the entire sky and essentially all levels of obscuration. This release provides a highly complete set of key measurements (emission line measurements and central velocity dispersions), with 99.9% measured redshifts and 98% black hole masses estimated (for unbeamed AGN outside the Galactic plane). The BASS DR2 AGN sample represents a unique census of nearby powerful AGN, spanning over 5 orders of magnitude in AGN bolometric luminosity, black hole mass, Eddington ratio, and obscuration. The public BASS DR2 sample and measurements can thus be used to answer fundamental questions about SMBH growth and its links to host galaxy evolution and feedback in the local universe, as well as open questions concerning SMBH physics. Here we provide a brief overview of the survey strategy, the key BASS DR2 measurements, data sets and catalogs, and scientific highlights from a series of DR2-based works.

astro-ph.GA

GaMPEN: A Machine Learning Framework for Estimating Bayesian Posteriors of Galaxy Morphological Parameters

We introduce a novel machine learning framework for estimating the Bayesian posteriors of morphological parameters for arbitrarily large numbers of galaxies. The Galaxy Morphology Posterior Estimation Network (GaMPEN) estimates values and uncertainties for a galaxy's bulge-to-total light ratio ($L_B/L_T$), effective radius ($R_e$), and flux ($F$). To estimate posteriors, GaMPEN uses the Monte Carlo Dropout technique and incorporates the full covariance matrix between the output parameters in its loss function. GaMPEN also uses a Spatial Transformer Network (STN) to automatically crop input galaxy frames to an optimal size before determining their morphology. This will allow it to be applied to new data without prior knowledge of galaxy size. Training and testing GaMPEN on galaxies simulated to match $z < 0.25$ galaxies in Hyper Suprime-Cam Wide $g$-band images, we demonstrate that GaMPEN achieves typical errors of $0.1$ in $L_B/L_T$, $0.17$ arcsec ($\sim 7\%$) in $R_e$, and $6.3\times10^4$ nJy ($\sim 1\%$) in $F$. GaMPEN's predicted uncertainties are well-calibrated and accurate ($<5\%$ deviation) -- for regions of the parameter space with high residuals, GaMPEN correctly predicts correspondingly large uncertainties. We also demonstrate that we can apply categorical labels (i.e., classifications such as "highly bulge-dominated") to predictions in regions with high residuals and verify that those labels are $\gtrsim 97\%$ accurate. To the best of our knowledge, GaMPEN is the first machine learning framework for determining joint posterior distributions of multiple morphological parameters and is also the first application of an STN to optical imaging in astronomy.

astro-ph.GA

BASS XXIV: The BASS DR2 Spectroscopic Line Measurements and AGN Demographics

We present the second catalog and data release of optical spectral line measurements and AGN demographics of the BAT AGN Spectroscopic Survey, which focuses on the of Swift-BAT hard X-ray detected AGNs. We use spectra from dedicated campaigns and publicly available archives to investigate spectral properties of most of the AGNs listed in the 70-month Swift-BAT all-sky catalog; specifically, 743 of the 746 unbeamed and unlensed AGNs (99.6%). We find a good correspondence between the optical emission line widths and the hydrogen column density distributions using the X-ray spectra, with a clear dichotomy of AGN types for NH = 10^22 cm-2. Based on optical emission-line diagnostics, we show that 48%-75% of BAT AGNs are classified as Seyfert, depending on the choice of emission lines used in the diagnostics. The fraction of objects with upper limits on line emission varies from 6% to 20%. Roughly 4% of the BAT AGNs have lines too weak to be placed on the most commonly used diagnostic diagram, [O III]λ5007/H\b{eta} versus [N II]λ6584/Hα, despite the high signal-to-noise ratio (S/N) of their spectra. This value increases to 35% in the [O III]λ5007/[O II]λ3727 diagram, owing to difficulties in line detection. Compared to optically-selected narrow-line AGNs in the Sloan Digital Sky Survey, the BAT narrow-line AGNs have a higher rate of reddening/extinction, with Hα/H\b{eta} > 5 (~ 36%), indicating that hard X-ray selection more effectively detects obscured AGNs from the underlying AGN population. Finally, we present a subpopulation of AGNs that feature complex broad-lines (34%, 250/743) or double-peaked narrow emission lines (2%, 17/743).

astro-ph.GA

BASS XXX: Distribution Functions of DR2 Eddington-ratios, Black Hole Masses, and X-ray Luminosities

We determine the low-redshift X-ray luminosity function (XLF), active black hole mass function (BHMF), and Eddington-ratio distribution function (ERDF) for both unobscured (Type 1) and obscured (Type 2) active galactic nuclei (AGN) using the unprecedented spectroscopic completeness of the BAT AGN Spectroscopic Survey (BASS) data release 2. In addition to a straightforward 1/Vmax approach, we also compute the intrinsic distributions, accounting for sample truncation by employing a forward modeling approach to recover the observed BHMF and ERDF. As previous BHMFs and ERDFs have been robustly determined only for samples of bright, broad-line (Type 1) AGNs and/or quasars, ours is the first directly observationally constrained BHMF and ERDF of Type 2 AGN. We find that after accounting for all observational biases, the intrinsic ERDF of Type 2 AGN is significantly skewed towards lower Eddington ratios than the intrinsic ERDF of Type 1 AGN. This result supports the radiation-regulated unification scenario, in which radiation pressure dictates the geometry of the dusty obscuring structure around an AGN. Calculating the ERDFs in two separate mass bins, we verify that the derived shape is consistent, validating the assumption that the ERDF (shape) is mass independent. We report the local AGN duty cycle as a function of mass and Eddington ratio, by comparing the BASS active BHMF with the local mass function for all SMBH. We also present the log N-log S of Swift-BAT 70-month sources.

astro-ph.HE

BAT AGN Spectroscopic Survey-XX: Molecular Gas in Nearby Hard X-ray Selected AGN Galaxies

We present the host galaxy molecular gas properties of a sample of 213 nearby (0.01 10^44 erg/s) increases by ~10-100 between a molecular gas mass of 10^8.7 Msun and 10^10.2 Msun. Higher Eddington ratio AGN galaxies tend to have higher molecular gas masses and gas fractions. Higher column density AGN galaxies (Log NH>23.4) are associated with lower depletion timescales and may prefer hosts with more gas centrally concentrated in the bulge that may be more prone to quenching than galaxy wide molecular gas. The significant average link of host galaxy molecular gas supply to SMBH growth may naturally lead to the general correlations found between SMBHs and their host galaxies, such as the correlations between SMBH mass and bulge properties and the redshift evolution of star formation and SMBH growth.

astro-ph.GA

Galaxy Morphology Network: A Convolutional Neural Network Used to Study Morphology and Quenching in $\sim 100,000$ SDSS and $\sim 20,000$ CANDELS Galaxies

We examine morphology-separated color-mass diagrams to study the quenching of star formation in $\sim 100,000$ ($z\sim0$) Sloan Digital Sky Survey (SDSS) and $\sim 20,000$ ($z\sim1$) Cosmic Assembly Near-Infrared Deep Extragalactic Legacy Survey (CANDELS) galaxies. To classify galaxies morphologically, we developed Galaxy Morphology Network (GaMorNet), a convolutional neural network that classifies galaxies according to their bulge-to-total light ratio. GaMorNet does not need a large training set of real data and can be applied to data sets with a range of signal-to-noise ratios and spatial resolutions. GaMorNet's source code as well as the trained models are made public as part of this work ( http://www.astro.yale.edu/aghosh/gamornet.html ). We first trained GaMorNet on simulations of galaxies with a bulge and a disk component and then transfer learned using $\sim25\%$ of each data set to achieve misclassification rates of $\lesssim5\%$. The misclassified sample of galaxies is dominated by small galaxies with low signal-to-noise ratios. Using the GaMorNet classifications, we find that bulge- and disk-dominated galaxies have distinct color-mass diagrams, in agreement with previous studies. For both SDSS and CANDELS galaxies, disk-dominated galaxies peak in the blue cloud, across a broad range of masses, consistent with the slow exhaustion of star-forming gas with no rapid quenching. A small population of red disks is found at high mass ($\sim14\%$ of disks at $z\sim0$ and $2\%$ of disks at $z \sim 1$). In contrast, bulge-dominated galaxies are mostly red, with much smaller numbers down toward the blue cloud, suggesting rapid quenching and fast evolution across the green valley. This inferred difference in quenching mechanism is in agreement with previous studies that used other morphology classification techniques on much smaller samples at $z\sim0$ and $z\sim1$.

astro-ph.GA

The BAT AGN Spectroscopic Survey -- XVIII. Searching for Supermassive Black Hole Binaries in the X-rays

Theory predicts that a supermassive black hole binary (SMBHB) could be observed as a luminous active galactic nucleus (AGN) that periodically varies on the order of its orbital timescale. In X-rays, periodic variations could be caused by mechanisms including relativistic Doppler boosting and shocks. Here we present the first systematic search for periodic AGNs using $941$ hard X-ray light curves (14-195 keV) from the first 105 months of the Swift Burst Alert Telescope (BAT) survey (2004-2013). We do not find evidence for periodic AGNs in Swift-BAT, including the previously reported SMBHB candidate MCG+11$-$11$-$032. We find that the null detection is consistent with the combination of the upper-limit binary population in AGNs in our adopted model, their expected periodic variability amplitudes, and the BAT survey characteristics. We have also investigated the detectability of SMBHBs against normal AGN X-ray variability in the context of the eROSITA survey. Under our assumptions of a binary population and the periodic signals they produce which have long periods of hundreds of days, up to $13$% true periodic binaries can be robustly distinguished from normal variable AGNs with the ideal uniform sampling. However, we demonstrate that realistic eROSITA sampling is likely to be insensitive to long-period binaries because longer observing gaps reduce their detectability. In contrast, large observing gaps do not diminish the prospect of detecting binaries of short, few-day periods, as 19% can be successfully recovered, the vast majority of which can be identified by the first half of the survey.

astro-ph.HE

Searching for Super-Eddington Quasars using a Photon Trapping Accretion Disc Model

Accretion onto black holes at rates above the Eddington limit has long been discussed in the context of supermassive black hole (SMBH) formation and evolution, providing a possible explanation for the presence of massive quasars at high redshifts (z$\gtrsim$7), as well as having implications for SMBH growth at later epochs. However, it is currently unclear whether such `super-Eddington' accretion occurs in SMBHs at all, how common it is, or whether every SMBH may experience it. In this work, we investigate the observational consequences of a simplistic model for super-Eddington accretion flows -- an optically thick, geometrically thin accretion disc (AD) where the inner-most parts experience severe photon-trapping, which is enhanced with increased accretion rate. The resulting spectral energy distributions (SEDs) show a dramatic lack of rest-frame UV, or even optical, photons. Using a grid of model SEDs spanning a wide range in parameter space (including SMBH mass and accretion rate), we find that large optical quasar surveys (such as SDSS) may be missing most of these luminous systems. We then propose a set of colour selection criteria across optical and infra-red colour spaces designed to select super-Eddington SEDs in both wide-field surveys (e.g., using SDSS, 2MASS and WISE) and deep & narrow-field surveys (e.g., COSMOS). The proposed selection criteria are a necessary first step in establishing the relevance of advection-affected super-Eddington accretion onto SMBHs at early cosmic epochs.

astro-ph.GA