SearcharxivSearch

arXiv subjects

Matthew R. Becker

Publications and source records attributed to Matthew R. Becker.

At least 19 recordsLinked to original sources

Diffhalos: A Generative Model of Cosmological Lightcones of Dark Matter Halos

We present a generative model of cosmological lightcones of dark matter halos, Diffhalos. In our model, we draw Monte Carlo samples of the halo mass function in a lightcone with a JAX-based implementation of the halo model, Halox, and we generate samples of subhalos by drawing from a model for the conditional subhalo mass function. We generate mass assembly histories (MAHs) using a normalizing flow trained on merger trees in cosmological N-body simulations. We show that Diffhalos can generate samples of halos, subhalos, and their MAHs with a statistical distribution that accurately approximates populations in simulated lightcones. As an example application, we use Diffhalos to calculate gradients of the halo and subhalo mass functions with respect to cosmological parameters. We conclude with a discussion of ongoing work using Diffhalos together with models of the galaxy--halo connection to make theoretical predictions for cosmological populations of galaxies, and to generate mock galaxy catalogs.

astro-ph.GA

Differentiable Forward Modeling for Efficient and Accurate Shear Inference

Forthcoming Stage-IV dark energy optical surveys, such as LSST, have the ambitious goal of measuring cosmological parameters at sub-percent precision. Realizing their full scientific potential requires very precise measurement of the cosmic shear signal and control of corresponding systematics. In this work, we present a modern implementation of the Bayesian shear inference framework in Schneider et al. (2015), in the case that the PSF and sky background are known. This framework automatically propagates the pixel-noise measurement error from each galaxy into the final shear estimate, and thus requires no external calibration to handle noise bias. As a first application of this new implementation, we infer the cosmic shear posterior from simulated images consisting of isolated exponential galaxies with semi-realistic noise levels. In this simplified scenario, we estimate the absolute multiplicative bias $|m|$ of our approach to be below $0.9 \times 10^{-3}\,[3\sigma]$ when the intrinsic distribution of galaxy properties is known, and below $1.3 \times 10^{-3}\,[3\sigma]$ when these distributions are inferred alongside shear. Additionally, we make progress towards the algorithm's computational feasibility in the context of modern wide-field surveys, where billions of galaxies must be processed, by leveraging differentiable forward models of galaxies, gradient-based samplers, and GPUs. Our final galaxy-fitting MCMC produces $300$ effective samples of galaxy properties in $0.45$ seconds per galaxy using a single A100 GPU. In the future, we seek to generalize our algorithm to handle selection, detection, and model shear biases so it can be applied to real survey data.

astro-ph.IM

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

astro-ph.IM

Uniform Rolling: An LSST Observing Cadence Offering Sufficient Survey Uniformity for Comprehensive Cosmological Analysis

The Legacy Survey of Space and Time (LSST) that will be carried out by the NSF-DOE Vera C. Rubin Observatory promises to be the defining survey of the next decade, supplying unprecedented access to the night sky to static science- and time-domain science-focused researchers alike. Maximizing the output of the broad remit of Rubin Observatory science requires a non-trivial survey strategy. For time-domain science, the most promising strategy designed so far is a rolling survey strategy, whereby a subset of the full LSST survey area is observed at higher rate compared with the nominal rate dictated by weather conditions and the observatory's technical constraints. This strategy is now the baseline approach for the LSST as a whole. Focusing on static science (galaxy clustering and weak lensing), we study how these time-domain-optimized rolling strategies affect the depth uniformity at intermediate years of the survey. We characterize the amount of survey area at high risk of being lost in static-science analyses of a baseline rolling LSST dataset due to an insufficient combination of survey contiguity and uniformity. At intermediate data releases, nearly half of the survey could be lost for static science, decreasing the Dark Energy figure of merit by approximately 40\%. We describe additional metrics focused on key analysis tasks, such as photometric redshifts and galaxy clustering. We propose a new strategy that returns the survey to uniformity at key release years, enabling use of the full survey area and restoring our metrics to the values they would have in a non-rolling cadence without loss of time domain data relative to a rolling survey with the same number of rolling cycles. This work has informed the third round of optimization of the survey strategy, and the new uniform rolling strategies have been incorporated into the baseline strategy.

astro-ph.CO

DiffstarPop: A generative physical model of galaxy star formation history

We present DiffstarPop, a differentiable forward model of cosmological populations of galaxy star formation histories (SFH). In the model, individual galaxy SFH is parametrized by Diffstar, which has parameters $\theta_{\rm SFH}$ that have a direct interpretation in terms of galaxy formation physics, such as star formation efficiency and quenching. DiffstarPop is a model for the statistical connection between $\theta_{\rm SFH}$ and the mass assembly history (MAH) of dark matter halos. We have formulated DiffstarPop to have the minimal flexibility needed to accurately reproduce the statistical distributions of galaxy SFH predicted by a diverse range of simulations, including the IllustrisTNG hydrodynamical simulation, the Galacticus semi-analytic model, and the UniverseMachine semi-empirical model. Our publicly available code written in JAX includes Monte Carlo generators that supply statistical samples of galaxy assembly histories that mimic the populations seen in each simulation, and can generate SFHs for $10^6$ galaxies in 1.1 CPU-seconds, or 0.03 GPU-seconds. We conclude the paper with a discussion of applications of DiffstarPop, which we are using to generate catalogs of synthetic galaxies populating the merger trees in cosmological N-body simulations.

astro-ph.GA

The little coadd that could: Estimating shear from coadded images

Upcoming wide field surveys will have many overlapping epochs of the same region of sky. The conventional wisdom is that in order to reduce the errors sufficiently for systematics-limited measurements, like weak lensing, we must do simultaneous fitting of all the epochs. Using current algorithms this will require a significant amount of computing time and effort. In this paper, we revisit the potential of using coadds for shear measurements. We show on a set of image simulations that the multiplicative shear bias can be constrained below the 0.1% level on coadds, which is sufficient for future lensing surveys. We see no significant differences between simultaneous fitting and coadded approaches for two independent shear codes: Metacalibration and BFD. One caveat of our approach is the assumption of a principled coadd, i.e. the PSF is mathematically well-defined for all the input images. This requires us to reject CCD images that do not fully cover the coadd region. We estimate that the number of epochs that must be rejected for a survey like LSST is on the order of 20%, resulting in a small loss in depth of less than 0.1 magnitudes. We also put forward a cell-based coaddition scheme that meets the above requirements for unbiased weak lensing shear estimation in the context of LSST.

astro-ph.CO

A Cohesive Deep Drilling Field Strategy for LSST Cosmology

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will image billions of astronomical objects in the wide-fast-deep primary survey and in a set of minisurveys including intensive observations of a group of deep drilling fields (DDFs). The DDFs are a critical piece of three key aspects of the LSST Dark Energy Science Collaboration (DESC) cosmological measurements: they provide a required calibration for photometric redshifts and weak gravitational lensing measurements and they directly contribute to cosmological constraints from the most distant type Ia supernovae. We present a set of cohesive DDF strategies fulfilling science requirements relevant to DESC and following the guidelines of the Survey Cadence Optimization Committee. We propose a method to estimate the observing strategy parameters and we perform simulations of the corresponding surveys. We define a set of metrics for each of the science case to assess the performance of the proposed observing strategies. We show that the most promising results are achieved with deep rolling surveys characterized by two sets of fields: ultradeep fields (z<1.1) observed at a high cadence with a large number of visits over a limited number of seasons; deep fields (z<0.7), observed with a cadence of ~3 nights for ten years. These encouraging results should be confirmed with realistic simulations using the LSST scheduler. A DDF budget of ~8.5% is required to design observing strategies satisfying all the cosmological requirements. A lower DDF budget lead to surveys that either do not fulfill photo-z/WL requirements or are not optimal for SNe Ia cosmology.

astro-ph.CO

Deep-field Metacalibration

We introduce deep-field metacalibration, a new technique that reduces the pixel noise in metacalibration estimators of weak lensing shear signals by using a deeper imaging survey for calibration. In standard metacalibration, when estimating the object's shear response, extra noise is added to correct the effect of shearing the noise in the image, increasing the uncertainty on shear estimates by ~ 20%. Our new deep-field metacalibration technique leverages a separate, deeper imaging survey to calculate calibrations with less degradation in image noise. We demonstrate that weak lensing shear measurement with deep-field metacalibration is unbiased up to second-order shear effects. We provide algorithms to apply this technique to imaging surveys and describe how to generalize it to shear estimators that rely explicitly on object detection (e.g., metacalibration). For the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), the improvement in weak lensing precision will depend on the somewhat unknown details of the LSST Deep Drilling Field (DDF) observations in terms of area and depth, the relative point-spread function properties of the DDF and main LSST surveys, and the relative contribution of pixel noise vs. intrinsic shape noise to the total shape noise in the survey. We conservatively estimate that the degradation in precision is reduced from 20% for metacalibration to ~ 5% or less for deep-field metacalibration, which we attribute primarily to the increased source density and reduced pixel noise contributions to the overall shape noise. Finally, we show that the technique is robust to sample variance in the LSST DDFs due to their large area, with the equivalent calibration error being ~ 0.1%. The deep-field metacalibration technique provides higher signal-to-noise weak lensing measurements while still meeting the stringent systematic error requirements of future surveys.

astro-ph.IM

Metadetection Weak Lensing for the Vera C. Rubin Observatory

Forthcoming astronomical imaging surveys will use weak gravitational lensing shear as a primary probe to study dark energy, with accuracy requirements at the 0.1% level. We present an implementation of the Metadetection shear measurement algorithm for use with the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST). This new code works with the data products produced by the LSST Science Pipelines, and uses the pipeline algorithms when possible. We tested the code using a new set of simulations designed to mimic LSST imaging data. The simulated images contained semi-realistic galaxies, stars with representative distributions of magnitudes and galactic spatial density, cosmic rays, bad CCD columns and spatially variable point spread functions. Bright stars were saturated and simulated ``bleed trails'' were drawn. Problem areas were interpolated, and the images were coadded into small cells, excluding images not fully covering the cell to guarantee a continuous point spread function. In all our tests the measured shear was accurate within the LSST requirements.

astro-ph.IM

Diffstar: A Fully Parametric Physical Model for Galaxy Assembly History

We present Diffstar, a smooth parametric model for the in-situ star formation history (SFH) of galaxies. Diffstar is distinct from conventional SFH models that are used to interpret the spectral energy distribution (SED) of an observed galaxy, because our model is parametrized directly in terms of basic features of galaxy formation physics. The Diffstar model assumes that star formation is fueled by the accretion of gas into the dark matter halo of the galaxy, and at the foundation of Diffstar is a parametric model for halo mass assembly, Diffmah. We include parametrized ingredients for the fraction of accreted gas that is eventually transformed into stars, $ε_{\rm ms},$ and for the timescale over which this transformation occurs, $τ_{\rm cons};$ some galaxies in Diffstar experience a quenching event at time $t_{\rm q},$ and may subsequently experience rejuvenated star formation. We fit the SFHs of galaxies predicted by the IllustrisTNG (TNG) and UniverseMachine (UM) simulations with the Diffstar parameterization, and show that our model is sufficiently flexible to describe the average stellar mass histories of galaxies in both simulations with an accuracy of $\sim0.1$ dex across most of cosmic time. We use Diffstar to compare TNG to UM in common physical terms, finding that: (i) star formation in UM is less efficient and burstier relative to TNG; (ii) galaxies in UM have longer gas consumption timescales, $τ_{\rm cons}$, relative to TNG; (iii) rejuvenated star formation is ubiquitous in UM, whereas quenched TNG galaxies rarely experience sustained rejuvenation; and (iv) in both simulations, the distributions of $ε_{\rm ms}$, $τ_{\rm cons}$, and $t_{\rm q}$ share a common characteristic dependence upon halo mass, and present significant correlations with halo assembly history. [Abridged]

astro-ph.GA

DSPS: Differentiable Stellar Population Synthesis

Models of stellar population synthesis (SPS) are the fundamental tool that relates the physical properties of a galaxy to its spectral energy distribution (SED). In this paper, we present DSPS: a python package for stellar population synthesis. All of the functionality in DSPS is implemented natively in the JAX library for automatic differentiation, and so our predictions for galaxy photometry are fully differentiable, and directly inherit the performance benefits of JAX, including portability onto GPUs. DSPS also implements several novel features, such as i) a flexible empirical model for stellar metallicity that incorporates correlations with stellar age, and ii) support for the diffstar model that provides a physically-motivated connection between the star formation history of a galaxy (SFH) and the mass assembly of its underlying dark matter halo. We detail a set of theoretical techniques for using autodiff to calculate gradients of predictions for galaxy SEDs with respect to SPS parameters that control a range of physical effects, including SFH, stellar metallicity, nebular emission, and dust attenuation. When forward modeling the colors of a synthetic galaxy population, we find that DSPS can provide a factor of 5 speedup over standard SPS codes on a CPU, and a factor of 300-400 on a modern GPU. When coupled with gradient-based techniques for optimization and inference, DSPS makes it practical to conduct expansive likelihood analyses of simulation-based models of the galaxy--halo connection that fully forward model galaxy spectra and photometry.

astro-ph.GA

A Differentiable Model of the Assembly of Individual and Populations of Dark Matter Halos

We present a new empirical model for the mass assembly of dark matter halos. We approximate the growth of individual halos as a simple power-law function of time, where the power-law index smoothly decreases as the halo transitions from the fast-accretion regime at early times, to the slow-accretion regime at late times. Using large samples of halo merger trees taken from high-resolution cosmological simulations, we demonstrate that our 3-parameter model, Diffmah, can approximate halo growth with a typical accuracy of 0.1 dex for t > 1 Gyr for all halos of present-day mass greater than 10^11Msun, including subhalos and host halos in gravity-only simulations, as well as in the TNG hydrodynamical simulation. We additionally present a new model for the assembly of halo populations, DiffmahPop, which not only reproduces average mass growth across time, but also faithfully captures the diversity with which halos assemble their mass. Our python implementation is based on the autodiff library JAX, and so our model self-consistently captures the mean and variance of halo mass accretion rate across cosmic time. We show that the connection between halo assembly and the large-scale density field, known as halo assembly bias, is accurately captured by Diffmah, and that residual errors in our approximations to halo assembly history exhibit a negligible residual correlation with the density field. Our publicly available source code can be used to generate Monte Carlo realizations of cosmologically representative halo histories; our differentiable implementation facilitates the incorporation of our model into existing analytical halo model frameworks.

astro-ph.CO

Differentiable Predictions for Large Scale Structure with SHAMNet

In simulation-based models of the galaxy-halo connection, theoretical predictions for galaxy clustering and lensing are typically made based on Monte Carlo realizations of a mock universe. In this paper, we use Subhalo Abundance Matching (SHAM) as a toy model to introduce an alternative to stochastic predictions based on mock population, demonstrating how to make simulation-based predictions for clustering and lensing that are both exact and differentiable with respect to the parameters of the model. Conventional implementations of SHAM are based on iterative algorithms such as Richardson-Lucy deconvolution; here we use the JAX library for automatic differentiation to train SHAMNet, a neural network that accurately approximates the stellar-to-halo mass relation (SMHM) defined by abundance matching. In our approach to making differentiable predictions for large scale structure, we map parameterized PDFs onto each simulated halo, and calculate gradients of summary statistics of the galaxy distribution by using autodiff to propagate the gradients of the SMHM through the statistical estimators used to measure one- and two-point functions. Our techniques are quite general, and we conclude with an overview of how they can be applied in tandem with more complex, higher-dimensional models, creating the capability to make differentiable predictions for the multi-wavelength universe of galaxies.

astro-ph.CO

ADDGALS: Simulated Sky Catalogs for Wide Field Galaxy Surveys

We present a method for creating simulated galaxy catalogs with realistic galaxy luminosities, broad-band colors, and projected clustering over large cosmic volumes. The technique, denoted ADDGALS (Adding Density Dependent GAlaxies to Lightcone Simulations), uses an empirical approach to place galaxies within lightcone outputs of cosmological simulations. It can be applied to significantly lower-resolution simulations than those required for commonly used methods such as halo occupation distributions, subhalo abundance matching, and semi-analytic models, while still accurately reproducing projected galaxy clustering statistics down to scales of r ~ 100 kpc/h. We show that \addgals\ catalogs reproduce several statistical properties of the galaxy distribution as measured by the Sloan Digital Sky Survey (SDSS) main galaxy sample, including galaxy number densities, observed magnitude and color distributions, as well as luminosity- and color-dependent clustering. We also compare to cluster-galaxy cross correlations, where we find significant discrepancies with measurements from SDSS that are likely linked to artificial subhalo disruption in the simulations. Applications of this model to simulations of deep wide-area photometric surveys, including modeling weak-lensing statistics, photometric redshifts, and galaxy cluster finding are presented in DeRose et al (2019), and an application to a full cosmology analysis of Dark Energy Survey (DES) Year 3 like data is presented in DeRose etl al (2021). We plan to publicly release a 10,313 square degree catalog constructed using ADDGALS with magnitudes appropriate for several existing and planned surveys, including SDSS, DES, VISTA, WISE, and LSST.

astro-ph.CO

Modeling Redshift-Space Clustering with Abundance Matching

We explore the degrees of freedom required to jointly fit projected and redshift-space clustering of galaxies selected in three bins of stellar mass from the Sloan Digital Sky Survey Main Galaxy Sample (SDSS MGS) using a subhalo abundance matching (SHAM) model. We employ emulators for relevant clustering statistics in order to facilitate our analysis, leading to large speed gains with minimal loss of accuracy. We are able to simultaneously fit the projected and redshift-space clustering of the two most massive galaxy samples that we consider with just two free parameters: scatter in stellar mass at fixed SHAM proxy and the dependence of the SHAM proxy on dark matter halo concentration. We find some evidence for models that include velocity bias, but including orphan galaxies improves our fits to the lower mass samples significantly. We also model the clustering signals of specific star formation rate (SSFR) selected samples using conditional abundance matching (CAM). We obtain acceptable fits to projected and redshift-space clustering as a function of SSFR and stellar mass using two CAM variants, although the fits are worse than for stellar mass-selected samples alone. By incorporating non-unity correlations between the CAM proxy and SSFR we are able to resolve previously identified discrepancies between CAM predictions and SDSS observations of the environmental dependence of quenching for isolated central galaxies.

astro-ph.CO

Mitigating Shear-dependent Object Detection Biases with Metacalibration

Metacalibration is a new technique for measuring weak gravitational lensing shear that is unbiased for isolated galaxy images. In this work we test metacalibration with overlapping, or ``blended'' galaxy images. Using standard metacalibration, we find a few percent shear measurement bias for galaxy densities relevant for current surveys, and that this bias increases with increasing galaxy number density. We show that this bias is not due to blending itself, but rather to shear-dependent object detection. If object detection is shear independent, no deblending of images is needed, in principle. We demonstrate that detection biases are accurately removed when including object detection in the metacalibration process, a technique we call metadetection. This process involves applying an artificial shear to images of small regions of sky and performing detection on the sheared images, as well as measurements that are used to calculate a shear response. We demonstrate that the method can accurately recover weak shear signals even in highly blended scenes. In the metacalibration process, the space between objects is sheared coherently, which does not perfectly match the real universe in which some, but not all, galaxy images are sheared coherently. We find that even for the worst case scenario, in which the space between objects is completely unsheared, the resulting shear bias is at most a few tenths of a percent for future surveys. We discuss additional technical challenges that must be met in order to implement metadetection for real surveys.

astro-ph.CO

A Redefinition of the Halo Boundary Leads to a Simple yet Accurate Halo Model of Large Scale Structure

We present a model for the halo--mass correlation function that explicitly incorporates halo exclusion. We assume that halos trace mass in a way that can be described using a single scale-independent bias parameter. However, our model exhibits scale dependent biasing due to the impact of halo-exclusion, the use of a ``soft'' (i.e. not infinitely sharp) halo boundary, and differences in the one halo term contributions to $ξ_{\rm hm}$ and $ξ_{\rm mm}$. These features naturally lead us to a redefinition of the halo boundary that lies at the ``by eye'' transition radius from the one--halo to the two--halo term in the halo--mass correlation function. When adopting our proposed definition, our model succeeds in describing the halo--mass correlation function with $\approx 2\%$ residuals over the radial range $0.1\ h^{-1}{\rm Mpc} < r < 80\ h^{-1}{\rm Mpc}$, and for halo masses in the range $10^{13}\ h^{-1}{\rm M}_{\odot} < M < 10^{15}\ h^{-1}{\rm M}_{\odot}$. Our proposed halo boundary is related to the splashback radius by a roughly constant multiplicative factor. Taking the 87-percentile as reference we find $r_{\rm t}/R_{\rm sp} \approx 1.3$. Surprisingly, our proposed definition results in halo abundances that are well described by the Press-Schechter mass function with $δ_{\rm sc}=1.449\pm 0.004$. The clustering bias parameter is offset from the standard background-split prediction by $\approx 10\%-15\%$. This level of agreement is comparable to that achieved with more standard halo definitions.

astro-ph.CO

The Aemulus Project IV: Emulating Halo Bias

Models of the spatial distribution of dark matter halos must achieve new levels of precision and accuracy in order to satisfy the requirements of upcoming experiments. In this work, we present a halo bias emulator for modeling the clustering of halos on large scales. It incorporates the cosmological dependence of the bias beyond the mapping of halo mass to peak height. The emulator makes substantial improvements in accuracy compared to the widely used Tinker et al. (2010) model. Halos in this work are defined using an overdensity criteria of 200 relative to the mean background density. Halo catalogs are produced for 40 N-body simulations as part of the Aemulus project at snapshots from z=3 to z=0. The emulator is trained over the mass range $6\times10^{12}-7\times10^{15}\ h^{-1}M_{\odot}$. Using an additional suite of 35 simulations, we determine that the precision of the emulator is redshift dependent, achieving sub-percent levels for a majority of the redshift range. Two additional simulation suites are used to test the ability of the emulator to extrapolate to higher and lower masses. Our high-resolution simulation suite is used to develop an extrapolation scheme in which the emulator asymptotes to the Tinker et al. (2010) model at low mass, achieving ~3% accuracy down to $10^{11}\ h^{-1}M_{\odot}$. Finally, we present a method to propagate emulator modeling uncertainty into an error budget. Our emulator is made publicly available at \url{https://github.com/AemulusProject/bias_emulator}.

astro-ph.CO