SearcharxivSearch

arXiv subjects

Adrian E. Bayer

Publications and source records attributed to Adrian E. Bayer.

At least 19 recordsLinked to original sources

Field-Level Baryon Acoustic Oscillation Reconstruction of the DESI DR1 Luminous Red Galaxies with Linear Field Transformer (LiFT)

We present the first application of neural field-level baryon-acoustic oscillation (BAO) reconstruction to real spectroscopic survey data. We develop Linear Field Transformer (LiFT), a 3D vision transformer that takes as input the observed galaxy field, its standard reconstruction, and a set of context channels encoding local line of sight, survey coverage, and redshift, and learns to correct standard reconstruction toward the linear density field. We construct a forward-modeling pipeline to produce mock lightcones similar to the DESI Data Release 1 (DR1) luminous red galaxy (LRG) sample, and train LiFT on these. We validate LiFT on held-out simulations, as well as on additional mocks which differ in gravity solver, halo finder, HOD, cosmology, and fiber-assignment history, as well as on mocks analyzed with a distorted distance-redshift relation; ultimately, we find unbiased dilation parameters with consistently tighter constraints than standard reconstruction. Applied to the DESI DR1 LRGs, LiFT improves errors on $α_{\rm iso}$ by 13%, 24%, and 30% and on $α_{\rm AP}$ by 5%, 22%, and 30% relative to the DESI DR1 standard reconstruction analysis in the LRG1, LRG2, and LRG3 bins respectively; meanwhile, our DR1 central values remain consistent with DR1 and DR2. This equates to a factor of 1.2, 1.7 and 2.0 increase in Figure of Merit (or effective survey volume) if one were to only use standard reconstruction. Ultimately, these results establish LiFT as a validated, survey-ready tool for current and upcoming galaxy surveys.

astro-ph.CO

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.

astro-ph.IM

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.

cs.CL

MujicΛ: Reconstructing Initial Conditions from Incomplete Redshift Surveys with Projected Optimization

In this paper, we introduce MujicΛ (Mapping the Universe with Jax-based Initial Condition ReconstrΛction), an optimization-based framework for reconstructing initial conditions from realistic galaxy spectroscopic redshift surveys. Unlike standard optimization-based approaches, MujicΛ augments the L-BFGS algorithm with a projection operator and rank-order matching to enforce Gaussianity of the initial conditions and substantially improve robustness to incomplete survey geometries. We validate MujicΛ on a mock lightcone catalog derived from semi-analytic models applied to the Millennium simulation. We construct a differentiable forward model that incorporates a fast particle-mesh simulation at megaparsec resolution and a comprehensive treatment of observational effects and survey incompleteness. MujicΛ reaches good agreement with the true density field down to the scale of the forward model, while maintaining consistency with the Gaussian prior through the projection step. It also broadly recovers the cosmic web classification, underscoring its value for deciphering environmental information in galaxy evolution studies. Beyond its key role in next-generation constrained simulations, the methodology offers a practical way to generate initial guesses and speed up field-level inference, especially for upcoming large-scale galaxy surveys.

astro-ph.CO

Field-Level Inference from Galaxies: BAO Reconstruction

Baryon acoustic oscillations (BAO) underpin the key cosmological results from modern spectroscopic galaxy surveys, but nonlinear gravitational evolution limits the precision achievable with traditional analysis methods. To overcome this, we develop field-level inference for BAO, first reconstructing the initial linear density field and then fitting the BAO signal therein. We benchmark three reconstruction methods: (i) traditional reconstruction based on the Zel'dovich approximation, (ii) explicit field-level inference using differentiable forward modeling with hybrid effective field theory, and (iii) implicit field-level inference using a convolutional neural network to augment traditional reconstruction. Using DESI-like Luminous Red Galaxy (LRG) and Bright Galaxy Survey (BGS) catalogs, we find that field-level approaches significantly sharpen the BAO feature relative to traditional reconstruction. For LRGs, explicit field-level inference improves constraints on the BAO scale parameters ($α_{\rm iso}, α_{\rm ap}$) by 26%, while implicit inference improves constraints by 35%, corresponding to a 2.4$\times$ improvement in figure of merit. For the higher-density, lower-redshift BGS sample, field-level inference enables information extraction from smaller scales, yielding an improvement in constraints of up to 46%, corresponding to a 3.2$\times$ improvement in figure of merit. Crucially, we address longstanding concerns regarding the robustness of field-level reconstruction by leveraging 1,000 mock realizations to perform extensive coverage tests. Our results are both unbiased and statistically well-calibrated, maintaining nominal coverage even when using tight simulation-informed priors and under model misspecification.

astro-ph.CO

Interpreting Cosmological Information from Neural Networks in the Hydrodynamic Universe

What happens when a black box (neural network) meets a black box (simulation of the Universe)? Recent work has shown that convolutional neural networks (CNNs) can infer cosmological parameters from the matter density field in the presence of complex baryonic processes. A key question that arises is, which parts of the cosmic web is the neural network obtaining information from? We shed light on the matter by identifying the Fourier scales, density scales, and morphological features of the cosmic web that CNNs pay most attention to. We find that CNNs extract cosmological information from both high and low density regions: overdense regions provide the most information per pixel, while underdense regions -- particularly deep voids and their surroundings -- contribute significantly due to their large spatial extent and coherent spatial features. Remarkably, we demonstrate that there is negligible degradation in cosmological constraining power after aggressive cutting in both maximum Fourier scale and density. Furthermore, we find similar results when considering both hydrodynamic and gravity-only simulations, implying that neural networks can marginalize over baryonic effects with minimal loss in cosmological constraining power. Our findings point to practical strategies for optimal and robust field-level cosmological inference in the presence of uncertainly modeled astrophysics.

astro-ph.CO

Robust CMB B-mode analysis with Needlet-ILC and simulation-based inference

We explore a novel analysis framework for parameter inference with large-scale CMB polarization data. Our method uses simulation-based inference combined with the needlet internal linear combination (NILC) algorithm and cross-correlation-based statistics to compress the data into a vector that is robust to model misspecification and small enough to be amenable to neural posterior estimation with normalizing flows. By leveraging this compressed data representation, our method enables the robust use of the anisotropic and non-Gaussian information in the foreground fields to more accurately separate the CMB polarization signal from these contaminants. Using an idealized ground-based experimental setup inspired by the Simons Observatory Small Aperture Telescopes, we demonstrate improved statistical constraining power for the tensor-to-scalar ratio $r$ compared to the (constrained) NILC algorithm and improved robustness to complex foregrounds compared to other techniques in the literature. Trained on a relatively simple semi-analytical foreground model, the method yields unbiased $r$ results across a range of PySM Galactic foreground simulations, including the high-complexity d12 model, for which we obtain $r=(1.09 \pm 0.27)\cdot 10^{-2}$ for input $r=0.01$ and sky fraction $f_{\mathrm{sky}} = 0.21$. We thus demonstrate the feasibility and advantages of a complete, maps-to-parameters, simulation-based analysis of large-scale CMB polarization for current ground-based observatories.

astro-ph.CO

jFoF: GPU Cluster Finding with Gradient Propagation

We present jFoF, a fully GPU-native Friends-of-Friends (FoF) halo finder designed for both high-performance simulation analysis and differentiable modeling. Implemented in JAX, jFoF achieves end-to-end acceleration by performing all neighbor searches, label propagation, and group construction directly on GPUs, eliminating costly host--device transfers. We introduce two complementary neighbor-search strategies, a standard k-d tree and a novel linked-cell grid, and demonstrate that jFoF attains up to an order-of-magnitude speedup compared to optimized CPU implementations while maintaining consistent halo catalogs. Beyond performance, jFoF enables gradient propagation through discrete halo-finding operations via both frozen-assignment and topological optimization modes. Using a topological optimization approach via a REINFORCE-style estimator, our approach allows smooth optimization of halo connectivity and membership, bridging continuous simulation fields with discrete structure catalogs. These capabilities make jFoF a foundation for differentiable inference, enabling end-to-end, gradient-based optimization of structure formation models within GPU-accelerated astrophysical pipelines. We make our code publicly available at https://github.com/bhorowitz/jFOF/.

astro-ph.IM

Impact of Simulation Box Size for Weak Lensing: Replication and Super-Sample Effects

We quantify the bias caused by small simulation box size on weak lensing observables and covariances, considering both replication and super-sample effects for a range of higher-order statistics. Using two simulation suites -- one comprising large boxes ($3750\,h^{-1}{\rm Mpc}$) and another constructed by tiling small boxes ($625\,h^{-1}{\rm Mpc}$) -- we generate full-sky convergence maps and extract $10^\circ \times 10^\circ$ patches via a Fibonacci grid. We consider biases in the mean and covariance of the angular power spectrum, bispectrum (up to $\ell=3000$), PDF, peak/minima counts, and Minkowski functionals. By first identifying lines of sight that are impacted by replications, we find that replication causes a O$(10\%)$ bias in the PDF and Minkowski functionals, and a O$(1\%)$ bias in other summary statistics. Replication also causes a O$(10\%)$ bias in the covariances, increasing with source redshift and $\ell$, reaching $\sim25\%$ for $z_s=2.5$. We additionally show that replication leads to heavy biases (up to O$(100\%)$ at high redshift) when performing gnomonic projection on a patch that is centered along a direction of replication. We then identify the lines of sight that are minimally affected by replication, and use the corresponding patches to isolate and study super-sample effects, finding that, while the mean values agree to within $1\%$, the variances differ by O$(10\%)$ for $z_s\leq2.5$. We show that these effects remain in the presence of noise and smoothing scales typical of the DES, KiDS, HSC, LSST, Euclid, and Roman surveys. We also discuss how these effects scale as a function of box size. Our results highlight the importance of large simulation volumes for accurate lensing statistics and covariance estimation.

astro-ph.CO

CosmoBench: A Multiscale, Multiview, Multitask Cosmology Benchmark for Geometric Deep Learning

Cosmological simulations provide a wealth of data in the form of point clouds and directed trees. A crucial goal is to extract insights from this data that shed light on the nature and composition of the Universe. In this paper we introduce CosmoBench, a benchmark dataset curated from state-of-the-art cosmological simulations whose runs required more than 41 million core-hours and generated over two petabytes of data. CosmoBench is the largest dataset of its kind: it contains 34 thousand point clouds from simulations of dark matter halos and galaxies at three different length scales, as well as 25 thousand directed trees that record the formation history of halos on two different time scales. The data in CosmoBench can be used for multiple tasks -- to predict cosmological parameters from point clouds and merger trees, to predict the velocities of individual halos and galaxies from their collective positions, and to reconstruct merger trees on finer time scales from those on coarser time scales. We provide several baselines on these tasks, some based on established approaches from cosmological modeling and others rooted in machine learning. For the latter, we study different approaches -- from simple linear models that are minimally constrained by symmetries to much larger and more computationally-demanding models in deep learning, such as graph neural networks. We find that least-squares fits with a handful of invariant features sometimes outperform deep architectures with many more parameters and far longer training time. Still there remains tremendous potential to improve these baselines by combining machine learning and cosmology to fully exploit the data. CosmoBench sets the stage for bridging cosmology and geometric deep learning at scale. We invite the community to push the frontier of scientific discovery by engaging with this dataset, available at https://cosmobench.streamlit.app

cs.LG

The Denario project: Deep knowledge AI agents for scientific discovery

We present Denario, an AI multi-agent system designed to serve as a scientific research assistant. Denario can perform many different tasks, such as generating ideas, checking the literature, developing research plans, writing and executing code, making plots, and drafting and reviewing a scientific paper. The system has a modular architecture, allowing it to handle specific tasks, such as generating an idea, or carrying out end-to-end scientific analysis using Cmbagent as a deep-research backend. In this work, we describe in detail Denario and its modules, and illustrate its capabilities by presenting multiple AI-generated papers generated by it in many different scientific disciplines such as astrophysics, biology, biophysics, biomedical informatics, chemistry, material science, mathematical physics, medicine, neuroscience and planetary science. Denario also excels at combining ideas from different disciplines, and we illustrate this by showing a paper that applies methods from quantum physics and machine learning to astrophysical data. We report the evaluations performed on these papers by domain experts, who provided both numerical scores and review-like feedback. We then highlight the strengths, weaknesses, and limitations of the current system. Finally, we discuss the ethical implications of AI-driven research and reflect on how such technology relates to the philosophy of science. We publicly release the code at https://github.com/AstroPilot-AI/Denario. A Denario demo can also be run directly on the web at https://huggingface.co/spaces/astropilot-ai/Denario, and the full app will be deployed on the cloud.

cs.AI

Transfer Learning Beyond the Standard Model

Machine learning enables powerful cosmological inference but typically requires many high-fidelity simulations covering many cosmological models. Transfer learning offers a way to reduce the simulation cost by reusing knowledge across models. We show that pre-training on the standard model of cosmology, $Λ$CDM, and fine-tuning on various beyond-$Λ$CDM scenarios -- including massive neutrinos, modified gravity, and primordial non-Gaussianities -- can enable inference with significantly fewer beyond-$Λ$CDM simulations. However, we also show that negative transfer can occur when strong physical degeneracies exist between $Λ$CDM and beyond-$Λ$CDM parameters. We consider various transfer architectures, finding that including bottleneck structures provides the best performance. Our findings illustrate the opportunities and pitfalls of foundation-model approaches in physics: pre-training can accelerate inference, but may also hinder learning new physics.

astro-ph.CO

The Power of the Cosmic Web

We study the cosmological information contained in the cosmic web, categorized as four structure types: nodes, filaments, walls, and voids, using the Quijote simulations and a modified nexus+ algorithm. We show that splitting the density field by the four structure types and combining the power spectrum in each provides much tighter constraints on cosmological parameters than using the power spectrum without splitting. We show the rich information contained in the cosmic web structures -- related to the Hessian of the density field -- for measuring all of the cosmological parameters, and in particular for constraining neutrino mass. We study the constraints as a function of Fourier scale, configuration space smoothing scale, and the underlying field. For the matter field with $k_{\rm max}=0.5\,h/{\rm Mpc}$, we find a factor of $\times20$ tighter constraints on neutrino mass when using smoothing scales larger than 12.5~Mpc/$h$, and $\times80$ tighter when using smoothing scales down to 1.95~Mpc/$h$. However, for the CDM+Baryon field we observe a more modest $\times1.7$ or $\times3.6$ improvement, for large and small smoothing scales respectively. We release our new python package for identifying cosmic structures pycosmmommf at https://github.com/James11222/pycosmommf to enable future studies of the cosmological information of the cosmic web.

astro-ph.CO

Field-Level Comparison and Robustness Analysis of Cosmological N-body Simulations

We present the first field-level comparison of cosmological N-body simulations, considering various widely used codes: Abacus, CUBEP$^3$M, Enzo, Gadget, Gizmo, PKDGrav, and Ramses. Unlike previous comparisons focused on summary statistics, we conduct a comprehensive field-level analysis: evaluating statistical similarity, quantifying implications for cosmological parameter inference, and identifying the regimes in which simulations are consistent. We begin with a traditional comparison using the power spectrum, cross-correlation coefficient, and visual inspection of the matter field. We follow this with a statistical out-of-distribution (OOD) analysis to quantify distributional differences between simulations, revealing insights not captured by the traditional metrics. We then perform field-level simulation-based inference (SBI) using convolutional neural networks (CNNs), training on one simulation and testing on others, including a full hydrodynamic simulation for comparison. We identify several causes of OOD behavior and biased inference, finding that resolution effects, such as those arising from adaptive mesh refinement (AMR), have a significant impact. Models trained on non-AMR simulations fail catastrophically when evaluated on AMR simulations, introducing larger biases than those from hydrodynamic effects. Differences in resolution, even when using the same N-body code, likewise lead to biased inference. We attribute these failures to a CNN's sensitivity to small-scale fluctuations, particularly in voids and filaments, and demonstrate that appropriate smoothing brings the simulations into statistical agreement. Our findings motivate the need for careful data filtering and the use of field-level OOD metrics, such as PQMass, to ensure robust inference.

astro-ph.CO

Simulation-Based Inference Benchmark for Weak Lensing Cosmology

Standard cosmological analysis, which relies on two-point statistics, fails to extract the full information of the data. This limits our ability to constrain with precision cosmological parameters. Thus, recent years have seen a paradigm shift from analytical likelihood-based to simulation-based inference. However, such methods require a large number of costly simulations. We focus on full-field inference, considered the optimal form of inference. Our objective is to benchmark several ways of conducting full-field inference to gain insight into the number of simulations required for each method. We make a distinction between explicit and implicit full-field inference. Moreover, as it is crucial for explicit full-field inference to use a differentiable forward model, we aim to discuss the advantages of having this property for the implicit approach. We use the sbi_lens package which provides a fast and differentiable log-normal forward model. This forward model enables us to compare explicit and implicit full-field inference with and without gradient. The former is achieved by sampling the forward model through the No U-Turns sampler. The latter starts by compressing the data into sufficient statistics and uses the Neural Likelihood Estimation algorithm and the one augmented with gradient. We perform a full-field analysis on LSST Y10 like weak lensing simulated mass maps. We show that explicit and implicit full-field inference yield consistent constraints. Explicit inference requires 630 000 simulations with our particular sampler corresponding to 400 independent samples. Implicit inference requires a maximum of 101 000 simulations split into 100 000 simulations to build sufficient statistics (this number is not fine tuned) and 1 000 simulations to perform inference. Additionally, we show that our way of exploiting the gradients does not significantly help implicit inference.

astro-ph.CO

Simulation-Efficient Cosmological Inference with Multi-Fidelity SBI

The simulation cost for cosmological simulation-based inference can be decreased by combining simulation sets of varying fidelity. We propose an approach to such multi-fidelity inference based on feature matching and knowledge distillation. Our method results in improved posterior quality, particularly for small simulation budgets and difficult inference problems.

astro-ph.CO

The HalfDome Multi-Survey Cosmological Simulations: N-body Simulations

Upcoming cosmological surveys have the potential to reach groundbreaking discoveries on multiple fronts, including the neutrino mass, dark energy, and inflation. Most of the key science goals require the joint analysis of datasets from multiple surveys to break parameter degeneracies and calibrate systematics. To realize such analyses, a large set of mock simulations that realistically model correlated observables is required. In this paper we present the N-body component of the HalfDome cosmological simulations, designed for the joint analysis of Stage-IV cosmological surveys, such as Rubin LSST, Euclid, SPHEREx, Roman, DESI, PFS, Simons Observatory, CMB-S4, and LiteBIRD. Our 300TB initial data release includes full-sky lightcones and halo catalogs between $z$=0--4 for 11 fixed cosmology realizations, as well as an additional run with local primordial non-Gaussianity ($f_{\rm NL}$=20). The simulations evolve $6144^3$ particles in a 3.75$\,h^{-1} {\rm Gpc}$ box, reaching a minimum halo mass of $\sim 6 \times 10^{12}\,h^{-1} M_\odot$ and maximum scale of $k \sim 1\,h{\rm Mpc}^{-1}$. Our data is publicly available: instructions to access the data and plans for future data releases can be found at https://halfdomesims.github.io.

astro-ph.CO

Massive $ν$s through the CNN lens: interpreting the field-level neutrino mass information in weak lensing

Modern cosmological surveys probe the Universe deep into the nonlinear regime, where massive neutrinos suppress cosmic structure. Traditional cosmological analyses, which use the 2-point correlation function to extract information, are no longer optimal in the nonlinear regime, and there is thus much interest in extracting beyond-2-point information to improve constraints on neutrino mass. Quantifying and interpreting the beyond-2-point information is thus a pressing task. We study the field-level information in weak lensing convergence maps using convolution neural networks. We find that the network performance increases as higher source redshifts and smaller scales are considered -- investigating up to a source redshift of 2.5 and $\ell_{\rm max}\simeq10^4$ -- verifying that massive neutrinos leave a distinct effect on weak lensing. However, the performance of the network significantly drops after scaling out the 2-point information from the maps, implying that most of the field-level information can be found in the 2-point correlation function alone. We quantify these findings in terms of the likelihood ratio and also use Integrated Gradient saliency maps to interpret which parts of the map the network is learning the most from. We find that, in the absence of noise, the network extracts a similar amount of information from the most overdense and underdense regions. However, upon adding noise, the information in underdense regions is distorted as noise disproportionately washes out void-like structures.

astro-ph.CO