SearcharxivSearch

arXiv subjects

Ludvig Doeser

Publications and source records attributed to Ludvig Doeser.

7 recordsLinked to original sources

Learning the Universe: Posterior Reliability of Neural Generative Models in High-Dimensional Field-Level Inference of Cosmic Initial Conditions

Accurate posterior estimation is central to scientific inference, as uncertainties determine what can be reliably learned from observational data. While Markov chain Monte Carlo methods provide asymptotic convergence guarantees, they are computationally demanding in high-dimensional settings. Neural network-based generative models for entire discretized 3D fields enable fast amortized inference but often lack convergence guarantees and principled accuracy assessment. Using Hamiltonian Monte Carlo to obtain reference posterior samples, we conduct a controlled field-level evaluation of an implicit generative model (Stochastic Interpolants) and an explicit likelihood-based model (GLOW normalizing flows). This comparison, unavailable in typical applications, enables the detection of posterior geometry failures that standard metrics cannot capture. As a case study, we consider the cosmological inverse problem of inferring cosmic initial conditions from present-day large-scale structure. To match the precision of modern cosmological data, this problem increasingly relies on complex, non-linear, and non-differentiable simulators, which are incompatible with gradient-based inference frameworks. Generative models offer a route to address these challenges, provided their inferred posteriors are reliable. In this work, we show that matching posterior means, marginal distributions, or achieving high cross-correlation does not imply correct uncertainty structure, as revealed by posterior variance fields and sample-based evaluations. Through this work, we aim to raise awareness of the challenges of uncertainty estimation in high-dimensional field-level settings, highlighting the importance of careful design and validation of neural generative approaches for scientific applications.

astro-ph.CO

The Manticore Project II: Bayesian digital twins of cosmic structure across the SDSS and BOSS volumes

We present Manticore-Deep, a high-resolution Bayesian field-level reconstruction of cosmic large-scale structure over a comoving volume of $(4~h^{-1}\mathrm{Gpc})^{3}$ to $z\approx0.7$ at ${\sim}4$~Mpc/h resolution. Extending the companion Manticore-Local analysis (Paper~I), Manticore-Deep jointly constrains five galaxy redshift surveys within a single hierarchical Bayesian framework using the BORG algorithm. The inference reconstructs primordial initial conditions evolved under gravity, yielding a posterior ensemble of three-dimensional density and velocity fields that causally reproduce the observed large-scale structure. A novel tiled inference strategy extends the reconstructed volume by more than an order of magnitude beyond Paper~I. Posterior realisations are consistent with LCDM, reproducing Gaussian isotropic initial conditions and the expected $z=0$ matter power spectrum, bispectrum, and halo mass function over the resolved scales. We validate the reconstruction using two independent template-free posterior-predictive tests against observations excluded from the inference. Cross-correlation with the \textit{Planck} PR3 CMB lensing map yields a cumulative detection significance of 7.4 $\sigma$, while velocity-weighted stacking of $64{,}750$ galaxy clusters on the \textit{Planck} 217~GHz map detects the kinetic Sunyaev--Zel'dovich effect at $3.5\sigma$, with a model-independent approach--recession split confirming the inferred velocities. Together, these tests validate both the projected-density and three-dimensional velocity fields recovered by Manticore-Deep. The BOSS Great Wall is recovered as a ${\sim}3\sigma$ overdensity consistent with LCDM across the posterior ensemble. Manticore-Deep establishes a benchmark for survey-depth constrained cosmological digital twins and reproducible field-level validation of large-scale structure reconstructions.

astro-ph.CO

Preparing for Rubin-LSST -- Detecting Brightest Cluster Galaxies with Machine Learning in the LSST DP0.2 simulation

The future Rubin Legacy Survey of Space and Time (LSST) is expected to deliver its first data release in the current of 2025. The upcoming survey will provide us with images of galaxy clusters in the optical to the near-infrared, with unrivalled coverage, depth and uniformity. The study of galaxy clusters informs us on the effect of environmental processes on galactic formation, which directly translates onto the formation of the brightest cluster galaxy (BCG). These massive galaxies present traces of the whole merger history of their host clusters, which can be in the shape of intra-cluster light (ICL) that surrounds them, tidal streams, or simply by the accumulated stellar mass that has been acquired over the past 10 billion years as they have cannibalized other galaxies in their surroundings. In an era where new data is being generated faster than humans can deal with, new methods involving machine learning have been emerging more and more in the most recent years. In the aim of preparing for the future LSST data release which will allow the observations of more than 20000 clusters and BCGs, we present in this paper different methods based on machine learning to detect these BCGs on LSST-like optical images. This study is done by making use of the simulated LSST Data Preview images. We find that the use of machine learning allows to accurately identify the BCG in up to 95% of clusters in our sample. Compared to more conventional red sequence extraction methods, the use of machine learning appears to be faster, more efficient and consistent, and does not require much, if any, pre-processing.

astro-ph.GA

Learning the Universe: $3\ h^{-1}{\rm Gpc}$ Tests of a Field Level $N$-body Simulation Emulator

We apply and test a field-level emulator for non-linear cosmic structure formation in a volume matching next-generation surveys. Inferring the cosmological parameters and initial conditions from which the particular galaxy distribution of our Universe was seeded can be achieved by comparing simulated data to observational data. Previous work has focused on building accelerated forward models that efficiently mimic these simulations. One of these accelerated forward models uses machine learning to apply a non-linear correction to the linear $z=0$ Zeldovich approximation (ZA) fields, closely matching the cosmological statistics in the $N$-body simulation. This emulator was trained and tested at $(h^{-1}{\rm Gpc})^3$ volumes, although cosmological inference requires significantly larger volumes. We test this emulator at $(3\ h^{-1}{\rm Gpc})^3$ by comparing emulator outputs to $N$-body simulations for eight unique cosmologies. We consider several summary statistics, applied to both the raw particle fields and the dark matter (DM) haloes. We find that the power spectrum, bispectrum and wavelet statistics of the raw particle fields agree with the $N$-body simulations within ${\sim} 5 \%$ at most scales. For the haloes, we find a similar agreement between the emulator and the $N$-body for power spectrum and bispectrum, though a comparison of the stacked profiles of haloes shows that the emulator has slight errors in the positions of particles in the highly non-linear interior of the halo. At these large $(3\ h^{-1}{\rm Gpc})^3$ volumes, the emulator can create $z=0$ particle fields in a thousandth of the time required for $N$-body simulations and will be a useful tool for large-scale cosmological inference. This is a Learning the Universe publication.

astro-ph.CO

Learning the Universe: Learning to Optimize Cosmic Initial Conditions with Non-Differentiable Structure Formation Models

Making the most of next-generation galaxy clustering surveys requires overcoming challenges in complex, non-linear modelling to access the significant amount of information at smaller cosmological scales. Field-level inference has provided a unique opportunity beyond summary statistics to use all of the information of the galaxy distribution. However, addressing current challenges often necessitates numerical modelling that incorporates non-differentiable components, hindering the use of efficient gradient-based inference methods. In this paper, we introduce Learning the Universe by Learning to Optimize (LULO), a gradient-free framework for reconstructing the 3D cosmic initial conditions. Our approach advances deep learning to train an optimization algorithm capable of fitting state-of-the-art non-differentiable simulators to data at the field level. Importantly, the neural optimizer solely acts as a search engine in an iterative scheme, always maintaining full physics simulations in the loop, ensuring scalability and reliability. We demonstrate the method by accurately reconstructing initial conditions from $M_{200\mathrm{c}}$ halos identified in a dark matter-only $N$-body simulation with a spherical overdensity algorithm. The derived dark matter and halo overdensity fields exhibit $\geq80\%$ cross-correlation with the ground truth into the non-linear regime $k \sim 1h$ Mpc$^{-1}$. Additional cosmological tests reveal accurate recovery of the power spectra, bispectra, halo mass function, and velocities. With this work, we demonstrate a promising path forward to non-linear field-level inference surpassing the requirement of a differentiable physics model.

astro-ph.CO

COmoving Computer Acceleration (COCA): $N$-body simulations in an emulated frame of reference

$N$-body simulations are computationally expensive, so machine-learning (ML)-based emulation techniques have emerged as a way to increase their speed. Although fast, surrogate models have limited trustworthiness due to potentially substantial emulation errors that current approaches cannot correct for. To alleviate this problem, we introduce COmoving Computer Acceleration (COCA), a hybrid framework interfacing ML with an $N$-body simulator. The correct physical equations of motion are solved in an emulated frame of reference, so that any emulation error is corrected by design. This approach corresponds to solving for the perturbation of particle trajectories around the machine-learnt solution, which is computationally cheaper than obtaining the full solution, yet is guaranteed to converge to the truth as one increases the number of force evaluations. Although applicable to any ML algorithm and $N$-body simulator, this approach is assessed in the particular case of particle-mesh cosmological simulations in a frame of reference predicted by a convolutional neural network, where the time dependence is encoded as an additional input parameter to the network. COCA efficiently reduces emulation errors in particle trajectories, requiring far fewer force evaluations than running the corresponding simulation without ML. We obtain accurate final density and velocity fields for a reduced computational budget. We demonstrate that this method shows robustness when applied to examples outside the range of the training data. When compared to the direct emulation of the Lagrangian displacement field using the same training resources, COCA's ability to correct emulation errors results in more accurate predictions. COCA makes $N$-body simulations cheaper by skipping unnecessary force evaluations, while still solving the correct equations of motion and correcting for emulation errors made by ML.

astro-ph.IM

Bayesian Inference of Initial Conditions from Non-Linear Cosmic Structures using Field-Level Emulators

Analysing next-generation cosmological data requires balancing accurate modeling of non-linear gravitational structure formation and computational demands. We propose a solution by introducing a machine learning-based field-level emulator, within the Hamiltonian Monte Carlo-based Bayesian Origin Reconstruction from Galaxies (BORG) inference algorithm. Built on a V-net neural network architecture, the emulator enhances the predictions by first-order Lagrangian perturbation theory to be accurately aligned with full $N$-body simulations while significantly reducing evaluation time. We test its incorporation in BORG for sampling cosmic initial conditions using mock data based on non-linear large-scale structures from $N$-body simulations and Gaussian noise. The method efficiently and accurately explores the high-dimensional parameter space of initial conditions, fully extracting the cross-correlation information of the data field binned at a resolution of $1.95h^{-1}$ Mpc. Percent-level agreement with the ground truth in the power spectrum and bispectrum is achieved up to the Nyquist frequency $k_\mathrm{N} \approx 2.79h \; \mathrm{Mpc}^{-1}$. Posterior resimulations - using the inferred initial conditions for $N$-body simulations - show that the recovery of information in the initial conditions is sufficient to accurately reproduce halo properties. In particular, we show highly accurate $M_{200\mathrm{c}}$ halo mass function and stacked density profiles of haloes in different mass bins $[0.853,16]\times 10^{14}M_{\odot}h^{-1}$. As all available cross-correlation information is extracted, we acknowledge that limitations in recovering the initial conditions stem from the noise level and data grid resolution. This is promising as it underscores the significance of accurate non-linear modeling, indicating the potential for extracting additional information at smaller scales.

astro-ph.CO