SearcharxivSearch

arXiv subjects

Boris Leistedt

Publications and source records attributed to Boris Leistedt.

At least 19 recordsLinked to original sources

pop-cosmos: Galaxy size evolution across structural and star-formation classifications in COSMOS-Web

Galaxy sizes are correlated with stellar mass and redshift, as characterised by size scaling relations. The inferred forms of these scaling relations are sensitive to how galaxies are classified -- either by their star formation activity (e.g. specific star-formation rate, sSFR) or by their morphology markers (e.g. bulge-to-total ratio, S\'{e}rsic index). We combine stellar mass and sSFR estimates from pop-cosmos (a generative model trained on COSMOS2020 Spitzer IRAC $\textit{Ch.1} <26$) with size and morphology measurements from COSMOS-Web, obtaining $99,369$ galaxies. By investigating the size-mass and the size-redshift relations, we show that: (i) the sSFR/morphology splits give quantitatively different slopes, intercepts, and intrinsic scatter behaviour; (ii) intrinsic scatter depends on structural morphology but not on sSFR, which constrains the galaxy-halo connection; (iii) the quiescent and bulge-dominated size-mass relations both show double-power law breaks, but at different pivot masses, indicating that quenching and structural transformation occur on different time-scales; (iv) the morphology-dependent trends are only recoverable from space-based imaging. Further, the quiescent pivot mass $M_{\ast} \sim 10^{10.7}~\mathrm{M}_{\odot}$ coincides with the mass scale at which AGN (infrared torus) bolometric luminosity fraction peaks in transitioning galaxies, while the bulge-dominated pivot mass $M_{\ast} \sim 10^{11.1}~\mathrm{M}_{\odot}$ coincides with the halo mass at which AGN-driven baryonic redistribution peaks, tracing the interval over which AGN feedback ramps from quenching onset to structural transformation.

astro-ph.GA

pop-cosmos: Disentangling galaxy properties from observables using data-driven approaches

The physical processes that shape a galaxy's spectrum are strongly degenerate in observations, obscuring which processes act independently. Leveraging the pop-cosmos generative galaxy population model, we investigate how many independent degrees of freedom the rest-frame optical SED contains. We use a $\beta$-variational autoencoder (VAE) to compress a 16-parameter stellar population synthesis (SPS) description into a disentangled latent representation interpreted through mutual information (MI). We find that five independent dimensions suffice, corresponding to stellar mass, recent star formation, dust, and two degrees of freedom in the ionization state of the gas. Stellar metallicity and stellar age are not among these primary drivers; their spectral effects are distributed across the others rather than independently encoded. By tying each dimension to specific spectral features, this decomposition breaks the star-formation--dust--metallicity degeneracies that limit broadband photometry, and recovers the physical conditions of the gas in typical star-forming galaxies more cleanly than the line-ratio diagnostics in standard use.

astro-ph.GA

pop-cosmos: Forward modeling KiDS-1000 redshift distributions using realistic galaxy populations

The accuracy of the cosmological constraints from Stage~IV galaxy surveys will be limited by how well the galaxy redshift distributions can be inferred. We have addressed this challenging problem for the Kilo-Degree Survey (KiDS) cosmic shear sample by developing a forward-modeling framework with two main ingredients: (1) the \texttt{pop-cosmos} generative model for the evolving galaxy population, calibrated on \textit{Spitzer} IRAC $\textit{Ch.\,1}<26$ galaxies from COSMOS2020; and (2) a data model for noise and selection, machine-learned from the SURFS-based KiDS-Legacy-Like Simulations (SKiLLS). Applying KiDS tomographic binning to our synthetic photometric data, we infer redshift distributions in each of five bins directly from the population and data models, bypassing the need for spectroscopic reweighting. Keeping the data model fixed, we compare results using two different galaxy population models: \texttt{pop-cosmos}; and \texttt{shark}, the semi-analytic galaxy formation model used in SKiLLS. In the first ($0.1<z<0.3$) and last ($0.9<z<1.2$) tomographic bins we find systematic differences in the mean redshifts of $\Delta z\sim0.05$-$0.1$, comparable to the reported uncertainties from spectroscopic reweighting methods. This work paves the way for accurate redshift distribution calibration for Stage~IV surveys directly through forward modeling, thus providing an independent cross-check on spectroscopic-based calibrations which avoids their selection biases and incompleteness. We will use the \texttt{pop-cosmos} redshift distributions in an upcoming full KiDS cosmology reanalysis.

astro-ph.CO

pop-cosmos: Redshifts and physical properties of KiDS-1000 galaxies

Principled Bayesian inference of galaxy properties has not previously been performed for wide-area weak lensing surveys with millions of sources. We address this gap by applying the pop-cosmos generative model to perform spectral energy distribution (SED) fitting for 4 million KiDS-1000 galaxies. Calibrated on deep COSMOS2020 photometric data, pop-cosmos specifies a physically-motivated prior over the galaxy population up to $z \simeq 6$ in stellar population synthesis (SPS) parameter space. Using the Speculator SPS emulator with GPU-accelerated MCMC sampling, we perform full posterior inference at 8.2 GPU seconds per galaxy, obtaining joint constraints on galaxy redshifts and physical properties. We validate photometric redshifts against $\sim\!185,\!000$ KiDS galaxies cross-matched to DESI DR1 spectroscopic samples, achieving low bias ($2\times10^{-3}$), scatter ($\sigma_{\mathrm{MAD}}=0.03$), and outlier fraction (3.2%) for the Bright Galaxy Survey, with comparable performance (bias $3\times10^{-2}$, $\sigma_{\mathrm{MAD}}=0.05$, 1.0% outliers) for luminous red galaxies (LRGs). Within the LRG sample, we identify massive, dusty, star-forming contaminants at $z \simeq 0.4$ satisfying standard colour selections for quenched populations. We infer trends in stellar mass, star formation, metallicity, and dust across five tomographic redshift bins consistent with established scaling relations. Using specific star formation rate constraints, we identify $\sim$7% of KiDS-1000 galaxies as quenched, versus 37% implied by conservative colour cuts. This enables the construction of weak lensing samples defined by physical properties while mitigating intrinsic alignment systematics and preserving statistical power. Our analysis validates pop-cosmos out-of-sample, establishing it as a scalable approach for galaxy evolution and cosmological analyses with photometric surveys.

astro-ph.GA

Uniform Rolling: An LSST Observing Cadence Offering Sufficient Survey Uniformity for Comprehensive Cosmological Analysis

The Legacy Survey of Space and Time (LSST) that will be carried out by the NSF-DOE Vera C. Rubin Observatory promises to be the defining survey of the next decade, supplying unprecedented access to the night sky to static science- and time-domain science-focused researchers alike. Maximizing the output of the broad remit of Rubin Observatory science requires a non-trivial survey strategy. For time-domain science, the most promising strategy designed so far is a rolling survey strategy, whereby a subset of the full LSST survey area is observed at higher rate compared with the nominal rate dictated by weather conditions and the observatory's technical constraints. This strategy is now the baseline approach for the LSST as a whole. Focusing on static science (galaxy clustering and weak lensing), we study how these time-domain-optimized rolling strategies affect the depth uniformity at intermediate years of the survey. We characterize the amount of survey area at high risk of being lost in static-science analyses of a baseline rolling LSST dataset due to an insufficient combination of survey contiguity and uniformity. At intermediate data releases, nearly half of the survey could be lost for static science, decreasing the Dark Energy figure of merit by approximately 40\%. We describe additional metrics focused on key analysis tasks, such as photometric redshifts and galaxy clustering. We propose a new strategy that returns the survey to uniformity at key release years, enabling use of the full survey area and restoring our metrics to the values they would have in a non-rolling cadence without loss of time domain data relative to a rolling survey with the same number of rolling cycles. This work has informed the third round of optimization of the survey strategy, and the new uniform rolling strategies have been incorporated into the baseline strategy.

astro-ph.CO

Systematics mitigation for catalogue-based angular power spectra

Recent work has developed a formalism for computing angular power spectra directly from catalogues containing field values at discrete positions on the sky, thereby circumventing the need to create pixelised maps of the fields, as well as avoiding aliasing and finite-resolution effects. We adapt this formalism to incorporate template deprojection for mitigating systematic biases in the measured angular power spectra. We also introduce an alternative method of mitigating the `deprojection bias' - the loss of modes induced by deprojection - employing simple simulations to compute a transfer function. We find that this approach performs at least as well as existing methods, and is relatively insensitive to how well one can guess the true power spectrum of the observed field, except at the largest scales ($\ell \lesssim 3$). Additionally, we develop exact expressions for the bias introduced by deprojection in the shot-noise component, which further improves the accuracy of this approach. We test our formalism on simulated datasets, demonstrating its applicability both to discretely sampled fields, and to the special case of galaxy clustering, with the survey selection function defined in terms of a random catalogue or as a continuous sky map. After removing the bias in the shot noise and correcting for the remaining mode loss using a transfer function, our formalism produces unbiased measurements of the angular power spectrum in all scenarios tested here. Finally, we apply our formalism to real data and show it produces results consistent with the standard map-based pseudo-$C_\ell$ formalism. We implement our method in the public code NaMaster.

astro-ph.CO

pop-cosmos: Star formation over 12 Gyr from generative modelling of a deep infrared-selected galaxy catalogue

We study star formation over 12 Gyr using pop-cosmos, a generative model trained on 26-band photometry of 420,000 COSMOS2020 galaxies (IRAC Ch.1 $<26$). The model learns distributions over 16 SPS parameters via score-based diffusion, matching observed colours and magnitudes. We compute the star formation rate density (SFRD) to $z=3.5$ by directly integrating individual galaxy SFRs. The SFRD peaks at $z=1.3\pm0.1$, with peak value $0.08\pm0.01$ M$_{\odot}$ yr$^{-1}$ Mpc$^{-3}$. We classify star-forming (SF) and quiescent (Q) galaxies using specific SFR $<10^{-11}$ yr$^{-1}$, comparing with $NUVrJ$ colour selection. The sSFR criterion yields up to 20% smaller quiescent fractions across $0 0.95$) during the most recent $\sim$300 Myr, then sharp decorrelation with earlier star-forming epochs, marking clear quenching transitions. Massive ($10<\log_{10}(M_*/$M$_{\odot})<11$) galaxies quench on a time-scale of $\sim1$ Gyr, with mass assembly concentrated in their first 3.5 Gyr. Finally, AGN activity (infrared luminosity) peaks as massive ($\sim10^{10.5}$ M$_\odot$) galaxies approach the transition between star-forming and quiescent states, declining sharply once quiescence is established. This provides evidence that AGN feedback operates in a critical regime during the $\sim1$ Gyr quenching transition.

astro-ph.GA

pop-cosmos: Insights from generative modeling of a deep, infrared-selected galaxy population

We present an extension of the pop-cosmos model for the evolving galaxy population up to redshift $z\sim6$. The model is trained on distributions of observed colors and magnitudes, from 26-band photometry of $\sim420,000$ galaxies in the COSMOS2020 catalog with Spitzer IRAC $\textit{Ch. 1}<26$. The generative model includes a flexible distribution over 16 stellar population synthesis (SPS) parameters, and a depth-dependent photometric uncertainty model, both represented using score-based diffusion models. We use the trained model to predict scaling relationships for the galaxy population, such as the stellar mass function, star-forming main sequence, and gas-phase and stellar metallicity vs. mass relations, demonstrating reasonable-to-excellent agreement with previously published results. We explore the connection between mid-infrared emission from active galactic nuclei (AGN) and star-formation rate, finding high AGN activity for galaxies above the star-forming main sequence at $1\lesssim z\lesssim 2$. Using the trained population model as a prior distribution, we perform inference of the redshifts and SPS parameters for 429,669 COSMOS2020 galaxies, including 39,588 with publicly available spectroscopic redshifts. The resulting redshift estimates exhibit minimal bias ($\text{median}[\Delta_z]=-8\times10^{-4}$), scatter ($\sigma_\text{MAD}=0.0132$), and outlier fraction ($6.19\%$) for the full $0<z<6$ spectroscopic compilation. These results establish that pop-cosmos can achieve the accuracy and realism needed to forward-model modern wide--deep surveys for Stage IV cosmology. We publicly release pop-cosmos software, mock galaxy catalogs, and COSMOS2020 redshift and SPS parameter posteriors.

astro-ph.GA

Imaging systematics induced by galaxy sub-sample fluctuation: new systematics at second order

Imaging systematics refers to the inhomogeneous distribution of a galaxy sample caused by varying observing conditions and astrophysical foregrounds. Current mitigation methods correct the galaxy density fluctuations caused by imaging systematics assuming that all galaxies in a sample have the same galaxy density fluctuations. Under this assumption, the corrected sample cannot perfectly recover the true correlation function. We name this effect sub-sample systematics. For a galaxy sample, even if its overall sample statistics (redshift distribution n(z), galaxy bias b(z)), are accurately measured, n(z), b(z) can still vary across the observed footprint. It makes the correlation function amplitude of galaxy clustering higher, while correlation functions for galaxy-galaxy lensing and cosmic shear do not have noticeable change. Such a combination could potentially degenerate with physical signals on small angular scales, such as the amplitude of galaxy clustering, the impact of neutrino mass on the matter power spectrum, etc. Sub-sample systematics cannot be corrected using imaging systematics mitigation approaches that rely on the cross-correlation signal between imaging systematics maps and the observed galaxy density field. In this paper, we derive formulated expressions of sub-sample systematics, demonstrating its fundamental difference with other imaging systematics. We also provide several toy models to visualize this effect. Finally, we discuss a potential method to estimate and mitigate sub-sample systematics by forward modeling its behavior using Synthetic Source Injection.

astro-ph.CO

Impact of redshift distribution uncertainties on Lyman-break galaxy cosmological parameter inference

A significant number of Lyman-break galaxies (LBGs) with redshifts 3 < z < 5 are expected to be observed by the upcoming Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST). This will enable us to probe the universe at higher redshifts than is currently possible with cosmological galaxy clustering and weak lensing surveys. However, accurate inference of cosmological parameters requires precise knowledge of the redshift distributions of selected galaxies, where the number of faint objects expected from LSST alone will make spectroscopic based methods of determining these distributions extremely challenging. To overcome this difficulty, it may be possible to leverage the information in the large volume of photometric data alone to precisely infer these distributions. This could be facilitated using forward models, where in this paper we use stellar population synthesis (SPS) to estimate uncertainties on LBG redshift distributions for a 10 year LSST (LSSTY10) survey. We characterise some of the modelling uncertainties inherent to SPS by introducing a flexible parameterisation of the galaxy population prior, informed by observations of the galaxy stellar mass function (GSMF) and cosmic star formation density (CSFRD). These uncertainties are subsequently marginalised over and propagated to cosmological constraints in a Fisher forecast. Assuming a known dust attenuation model for LBGs, we forecast constraints on the sigma8 parameter comparable to Planck cosmic microwave background (CMB) constraints.

astro-ph.CO

Quantifying the Impact of LSST $u$-band Survey Strategy on Photometric Redshift Estimation and the Detection of Lyman-break Galaxies

The Vera C. Rubin Observatory will conduct the Legacy Survey of Space and Time (LSST), promising to discover billions of galaxies out to redshift 7, using six photometric bands ($ugrizy$) spanning the near-ultraviolet to the near-infrared. The exact number of and quality of information about these galaxies will depend on survey depth in these six bands, which in turn depends on the LSST survey strategy: i.e., how often and how long to expose in each band. $u$-band depth is especially important for photometric redshift (photo-$z$) estimation and for detection of high-redshift Lyman-break galaxies (LBGs). In this paper we use a simulated galaxy catalog and an analytic model for the LBG population to study how recent updates and proposed changes to Rubin's $u$-band throughput and LSST survey strategy impact photo-$z$ accuracy and LBG detection. We find that proposed variations in $u$-band strategy have a small impact on photo-$z$ accuracy for $z < 1.5$ galaxies, but the outlier fraction, scatter, and bias for higher redshift galaxies varies by up to 50%, depending on the survey strategy considered. The number of $u$-band dropout LBGs at $z \sim 3$ is also highly sensitive to the $u$-band depth, varying by up to 500%, while the number of $griz$-band dropouts is only modestly affected. Under the new $u$-band strategy recommended by the Rubin Survey Cadence Optimization Committee, we predict $u$-band dropout number densities of $110$ deg$^{-2}$ (3200 deg$^{-2}$) in year 1 (10) of LSST. We discuss the implications of these results for LSST cosmology.

astro-ph.CO

Galaxy Clustering with LSST: Effects of Number Count Bias from Blending

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will survey the southern sky to create the largest galaxy catalog to date, and its statistical power demands an improved understanding of systematic effects such as source overlaps, also known as blending. In this work we study how blending introduces a bias in the number counts of galaxies (instead of the flux and colors), and how it propagates into galaxy clustering statistics. We use the $300\,$deg$^2$ DC2 image simulation and its resulting galaxy catalog (LSST Dark Energy Science Collaboration et al. 2021) to carry out this study. We find that, for a LSST Year 1 (Y1)-like cosmological analyses, the number count bias due to blending leads to small but statistically significant differences in mean redshift measurements when comparing an observed sample to an unblended calibration sample. In the two-point correlation function, blending causes differences greater than 3$\sigma$ on scales below approximately $10'$, but large scales are unaffected. We fit $\Omega_{\rm m}$ and linear galaxy bias in a Bayesian cosmological analysis and find that the recovered parameters from this limited area sample, with the LSST Y1 scale cuts, are largely unaffected by blending. Our main results hold when considering photometric redshift and a LSST Year 5 (Y5)-like sample.

astro-ph.CO

An automated method for finding the most distant quasars

Upcoming surveys such as Euclid, the Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) and the Nancy Grace Roman Telescope (Roman) will detect hundreds of high-redshift (z > 7) quasars, but distinguishing them from the billions of other sources in these catalogues represents a significant data analysis challenge. We address this problem by extending existing selection methods by using both i) Bayesian model comparison on measured fluxes and ii) a likelihood-based goodness-of-fit test on images, which are then combined using the F_beta statistic (where beta is a parameter which can be tuned to prioritise completeness). The result is an automated, reproduceable and objective high-redshift quasar selection pipeline. We test this on both simulations and real data from the cross-matched Sloan Digital Sky Survey (SDSS) and UKIRT Infrared Deep Sky Survey (UKIDSS) catalogues. On this cross-matched dataset we achieve an area under the curve (AUC) score of up to 0.81 and an F_3 score of up to 0.79; or, if the completeness is fixed to be 0.9, then we can obtain an efficiency of 0.15. This is sufficient to be applied to the Euclid, LSST and Roman data when available.

astro-ph.IM

pop-cosmos: Scaleable inference of galaxy properties and redshifts with a data-driven population model

We present an efficient Bayesian method for estimating individual photometric redshifts and galaxy properties under a pre-trained population model (pop-cosmos) that was calibrated using purely photometric data. This model specifies a prior distribution over 16 stellar population synthesis (SPS) parameters using a score-based diffusion model, and includes a data model with detailed treatment of nebular emission. We use a GPU-accelerated affine invariant ensemble sampler to achieve fast posterior sampling under this model for 292,300 individual galaxies in the COSMOS2020 catalog, leveraging a neural network emulator (Speculator) to speed up the SPS calculations. We apply both the pop-cosmos population model and a baseline prior inspired by Prospector-$\alpha$, and compare these results to published COSMOS2020 redshift estimates from the widely-used EAZY and LePhare codes. For the $\sim 12,000$ galaxies with spectroscopic redshifts, we find that pop-cosmos yields redshift estimates that have minimal bias ($\sim10^{-4}$), high accuracy ($\sigma_\text{MAD}=7\times10^{-3}$), and a low outlier rate ($1.6\%$). We show that the pop-cosmos population model generalizes well to galaxies fainter than its $r<25$ mag training set. The sample we have analyzed is $\gtrsim3\times$ larger than has previously been possible via posterior sampling with a full SPS model, with average throughput of 15 GPU-sec per galaxy under the pop-cosmos prior, and 0.6 GPU-sec per galaxy under the Prospector prior. This paves the way for principled modeling of the huge catalogs expected from upcoming Stage IV galaxy surveys.

astro-ph.CO

pop-cosmos: A comprehensive picture of the galaxy population from COSMOS data

We present pop-cosmos: a comprehensive model characterizing the galaxy population, calibrated to $140,938$ ($r<25$ selected) galaxies from the Cosmic Evolution Survey (COSMOS) with photometry in $26$ bands from the ultra-violet to the infra-red. We construct a detailed forward model for the COSMOS data, comprising: a population model describing the joint distribution of galaxy characteristics and its evolution (parameterized by a flexible score-based diffusion model); a state-of-the-art stellar population synthesis (SPS) model connecting galaxies' instrinsic properties to their photometry; and a data-model for the observation, calibration and selection processes. By minimizing the optimal transport distance between synthetic and real data we are able to jointly fit the population- and data-models, leading to robustly calibrated population-level inferences that account for parameter degeneracies, photometric noise and calibration, and selection. We present a number of key predictions from our model of interest for cosmology and galaxy evolution, including the mass function and redshift distribution; the mass-metallicity-redshift and fundamental metallicity relations; the star-forming sequence; the relation between dust attenuation and stellar mass, star formation rate and attenuation-law index; and the relation between gas-ionization and star formation. Our model encodes a comprehensive picture of galaxy evolution that faithfully predicts galaxy colors across a broad redshift ($z<4$) and wavelength range.

astro-ph.GA

Data-Space Validation of High-Dimensional Models by Comparing Sample Quantiles

We present a simple method for assessing the predictive performance of high-dimensional models directly in data space when only samples are available. Our approach is to compare the quantiles of observables predicted by a model to those of the observables themselves. In cases where the dimensionality of the observables is large (e.g. multiband galaxy photometry), we advocate that the comparison is made after projection onto a set of principal axes to reduce the dimensionality. We demonstrate our method on a series of two-dimensional examples. We then apply it to results from a state-of-the-art generative model for galaxy photometry (pop-cosmos; arXiv:2402.00935) that generates predictions of colors and magnitudes by forward simulating from a 16-dimensional distribution of physical parameters represented by a score-based diffusion model. We validate the predictive performance of this model directly in a space of nine broadband colors. Although motivated by this specific example, we expect that the techniques we present will be broadly useful for evaluating the performance of flexible, non-parametric population models of this kind, and other settings where two sets of samples are to be compared.

astro-ph.IM

Hierarchical Bayesian inference of photometric redshifts with stellar population synthesis models

We present a Bayesian hierarchical framework to analyze photometric galaxy survey data with stellar population synthesis (SPS) models. Our method couples robust modeling of spectral energy distributions with a population model and a noise model to characterize the statistical properties of the galaxy populations and real observations, respectively. By self-consistently inferring all model parameters, from high-level hyper-parameters to SPS parameters of individual galaxies, one can separate sources of bias and uncertainty in the data.We demonstrate the strengths and flexibility of this approach by deriving accurate photometric redshifts for a sample of spectroscopically-confirmed galaxies in the COSMOS field, all with 26-band photometry and spectroscopic redshifts. We achieve a performance competitive with publicly-released photometric redshift catalogs based on the same data. Prior to this work, this approach was computationally intractable in practice due to the heavy computational load of SPS model calls; we overcome this challenge using with neural emulators. We find that the largest photometric residuals are associated with poor calibration for emission line luminosities and thus build a framework to mitigate these effects. This combination of physics-based modeling accelerated with machine learning paves the path towards meeting the stringent requirements on the accuracy of photometric redshift estimation imposed by upcoming cosmological surveys. The approach also has the potential to create new links between cosmology and galaxy evolution through the analysis of photometric datasets.

astro-ph.IM

Spurious correlations between galaxies and multi-epoch image stacks in the DESI Legacy Surveys

A non-negligible source of systematic bias in cosmological analyses of galaxy surveys is the on-sky modulation caused by foregrounds and variable image characteristics such as observing conditions. Standard mitigation techniques perform a regression between the observed galaxy density field and sky maps of the potential contaminants. Such maps are ad-hoc, lossy summaries of the heterogeneous sets of co-added exposures that contribute to the survey. We present a methodology to address this limitation, and extract the spurious correlations between the observed distribution of galaxies and arbitrary stacks of single-epoch exposures. We study four types of galaxies (LRGs, ELGs, QSOs, LBGs) in the three regions of the DESI Legacy Surveys (North, South, DES), which results in twelve samples with varying levels and type of contamination. We find that the new technique outperforms the traditional ones in all cases, and is able to remove higher levels of contamination. This paves the way for new methods that extract more information from multi-epoch galaxy survey data and mitigate large-scale biases more effectively.

astro-ph.CO