SearcharxivSearch

arXiv subjects

Uros Seljak

Publications and source records attributed to Uros Seljak.

At least 19 recordsLinked to original sources

Can Microcanonical Langevin Dynamics Leverage Mini-Batch Gradient Noise?

Scaling inference methods such as Markov chain Monte Carlo to high-dimensional models remains a central challenge in Bayesian deep learning. A promising recent proposal, microcanonical Langevin Monte Carlo, has shown state-of-the-art performance across a wide range of problems. However, its reliance on full-dataset gradients makes it prohibitively expensive for large-scale problems. This paper addresses a fundamental question: Can microcanonical dynamics effectively leverage mini-batch gradient noise? We provide the first systematic study of this problem, establishing a novel continuous-time theoretical analysis of stochastic-gradient microcanonical dynamics. We reveal two critical failure modes: a theoretically derived bias due to anisotropic gradient noise and numerical instabilities in complex high-dimensional posteriors. To tackle these issues, we propose a principled gradient noise preconditioning scheme shown to significantly reduce this bias and develop a novel, energy-variance-based adaptive tuner that automates step size selection and dynamically informs numerical guardrails. The resulting algorithm is a robust and scalable microcanonical Monte Carlo sampler that achieves state-of-the-art performance on challenging high-dimensional inference tasks like Bayesian neural networks. Combined with recent ensemble techniques, our work unlocks a new class of stochastic microcanonical Langevin ensemble (SMILE) samplers for large-scale Bayesian inference.

cs.LG

Field-Level Inference from Galaxies: BAO Reconstruction

Baryon acoustic oscillations (BAO) underpin the key cosmological results from modern spectroscopic galaxy surveys, but nonlinear gravitational evolution limits the precision achievable with traditional analysis methods. To overcome this, we develop field-level inference for BAO, first reconstructing the initial linear density field and then fitting the BAO signal therein. We benchmark three reconstruction methods: (i) traditional reconstruction based on the Zel'dovich approximation, (ii) explicit field-level inference using differentiable forward modeling with hybrid effective field theory, and (iii) implicit field-level inference using a convolutional neural network to augment traditional reconstruction. Using DESI-like Luminous Red Galaxy (LRG) and Bright Galaxy Survey (BGS) catalogs, we find that field-level approaches significantly sharpen the BAO feature relative to traditional reconstruction. For LRGs, explicit field-level inference improves constraints on the BAO scale parameters ($α_{\rm iso}, α_{\rm ap}$) by 26%, while implicit inference improves constraints by 35%, corresponding to a 2.4$\times$ improvement in figure of merit. For the higher-density, lower-redshift BGS sample, field-level inference enables information extraction from smaller scales, yielding an improvement in constraints of up to 46%, corresponding to a 3.2$\times$ improvement in figure of merit. Crucially, we address longstanding concerns regarding the robustness of field-level reconstruction by leveraging 1,000 mock realizations to perform extensive coverage tests. Our results are both unbiased and statistically well-calibrated, maintaining nominal coverage even when using tight simulation-informed priors and under model misspecification.

astro-ph.CO

The Future of Artificial Intelligence and the Mathematical and Physical Sciences (AI+MPS)

This community paper developed out of the NSF Workshop on the Future of Artificial Intelligence (AI) and the Mathematical and Physics Sciences (MPS), which was held in March 2025 with the goal of understanding how the MPS domains (Astronomy, Chemistry, Materials Research, Mathematical Sciences, and Physics) can best capitalize on, and contribute to, the future of AI. We present here a summary and snapshot of the MPS community's perspective, as of Spring/Summer 2025, in a rapidly developing field. The link between AI and MPS is becoming increasingly inextricable; now is a crucial moment to strengthen the link between AI and Science by pursuing a strategy that proactively and thoughtfully leverages the potential of AI for scientific discovery and optimizes opportunities to impact the development of AI by applying concepts from fundamental science. To achieve this, we propose activities and strategic priorities that: (1) enable AI+MPS research in both directions; (2) build up an interdisciplinary community of AI+MPS researchers; and (3) foster education and workforce development in AI for MPS researchers and students. We conclude with a summary of suggested priorities for funding agencies, educational institutions, and individual researchers to help position the MPS community to be a leader in, and take full advantage of, the transformative potential of AI+MPS.

cs.AI

Machine Learning

This chapter gives an overview of the core concepts of machine learning (ML) -- the use of algorithms that learn from data, identify patterns, and make predictions or decisions without being explicitly programmed -- that are relevant to particle physics with some examples of applications to the energy, intensity, cosmic, and accelerator frontiers.

physics.data-an

Detecting Modeling Bias with Continuous Time Flow Models on Weak Lensing Maps

Simulation-based inference provides a powerful framework for extracting rich information from nonlinear scales in current and upcoming cosmological surveys, and ensuring its robustness requires stringent validation of forward models. In this work, we recast forward model validation as an out-of-distribution (OoD) detection problem within the framework of machine learning (ML)-based simulation-based inference (SBI). We employ probability density as the metric for OoD detection, and compare various density estimation techniques, demonstrating that field-level probability density estimation via continuous time flow models (CTFM) significantly outperforms feature-level approaches that combine scattering transform (ST) or convolutional neural networks (CNN) with normalizing flows (NFs), as well as NF-based field-level estimators, as quantified by the area under the receiver operating characteristic curve (AUROC). Our analysis shows that CTFM not only excels in detecting OoD samples but also provides a robust metric for model selection. Additionally, we verified CTFM maintains consistent efficacy across different cosmologies while mitigating the inductive biases inherent in NF architectures. Although our proof-of-concept study employs simplified forward modeling and noise settings, our framework establishes a promising pathway for identifying unknown systematics in the cosmology datasets.

astro-ph.CO

The Spectroscopic Stage-5 Experiment

The existence, properties, and dynamics of the dark sectors of our universe pose fundamental challenges to our current model of physics, and large-scale astronomical surveys may be our only hope to unravel these long-standing mysteries. In this white paper, we describe the science motivation, instrumentation, and survey plan for the next-generation spectroscopic observatory, the Stage-5 Spectroscopic Experiment (Spec-S5). Spec-S5 is a new all-sky spectroscopic instrument optimized to efficiently carry out cosmological surveys of unprecedented scale and precision. The baseline plan for Spec-S5 involves upgrading two existing 4-m telescopes to new 6-m wide-field facilities, each with a highly multiplexed spectroscopic instrument capable of simultaneously measuring the spectra of 13,000 astronomical targets. Spec-S5, which builds and improves on the hardware used for previous cosmology experiments, represents a cost-effective and rapid approach to realizing a more than 10$\times$ gain in spectroscopic capability compared to the current state-of-the-art represented by the Dark Energy Spectroscopic Instrument project (DESI). Spec-S5 will provide a critical scientific capability in the post-Rubin and post-DESI era for advancing cosmology, fundamental physics, and astrophysics in the 2030s.

astro-ph.CO

Initial Conditions from Galaxies: Machine-Learning Subgrid Correction to Standard Reconstruction

We present a hybrid method for reconstructing the primordial density from late-time halos and galaxies. Our approach involves two steps: (1) apply standard Baryon Acoustic Oscillation (BAO) reconstruction to recover the large-scale features in the primordial density field and (2) train a deep learning model to learn small-scale corrections on partitioned subgrids of the full volume. At inference, this correction is then convolved across the full survey volume, enabling scaling to large survey volumes. We train our method on both mock halo catalogs and mock galaxy catalogs in both configuration and redshift space from the Quijote $1(h^{-1}\,\mathrm{Gpc})^3$ simulation suite. When evaluated on held-out simulations, our combined approach significantly improves the reconstruction cross-correlation coefficient with the true initial density field and remains robust to moderate model misspecification. Additionally, we show that models trained on $1(h^{-1}\,\mathrm{Gpc})^3$ can be applied to larger boxes--e.g., $(3h^{-1}\,\mathrm{Gpc})^3$--without retraining. Finally, we perform a Fisher analysis on our method's recovery of the BAO peak, and find that it significantly improves the error on the acoustic scale relative to standard BAO reconstruction. Ultimately, this method robustly captures nonlinearities and bias without sacrificing large-scale accuracy, and its flexibility to handle arbitrarily large volumes without escalating computational requirements makes it especially promising for large-volume surveys like DESI.

astro-ph.CO

Local Primordial non-Gaussian Bias from Time Evolution

Primordial non-Gaussianity (PNG) is a signature of fundamental physics in the early universe that is probed by cosmological observations. It is well known that the local type of PNG generates a strong signal in the two-point function of large-scale structure tracers, such as galaxies. This signal, often termed ``scale-dependent bias'' is a generic feature of modulation of gravitational structure formation by a large-scale mode. It is less well-appreciated that the coefficient controlling this signal, $b_ϕ$, is closely connected to the time evolution of the tracer number density. This correspondence between time evolution and local PNG can be simply explained for a universal tracer whose mass function only depends on peak height, and more generally for non-universal tracers in the separate universe picture, which we validate in simulations. We also describe how to recover the bias of tracers subject to a survey selection function, and perform a simple demonstration on simulated galaxies. Since the local PNG amplitude in $n-$point statistics ($f_{\rm NL}$) is largely degenerate with the coefficient $b_ϕ$, this proof of concept study demonstrates that galaxy survey data can allow for more optimal and robust extraction of local PNG information from upcoming surveys.

astro-ph.CO

Microcanonical Langevin Ensembles: Advancing the Sampling of Bayesian Neural Networks

Despite recent advances, sampling-based inference for Bayesian Neural Networks (BNNs) remains a significant challenge in probabilistic deep learning. While sampling-based approaches do not require a variational distribution assumption, current state-of-the-art samplers still struggle to navigate the complex and highly multimodal posteriors of BNNs. As a consequence, sampling still requires considerably longer inference times than non-Bayesian methods even for small neural networks, despite recent advances in making software implementations more efficient. Besides the difficulty of finding high-probability regions, the time until samplers provide sufficient exploration of these areas remains unpredictable. To tackle these challenges, we introduce an ensembling approach that leverages strategies from optimization and a recently proposed sampler called Microcanonical Langevin Monte Carlo (MCLMC) for efficient, robust and predictable sampling performance. Compared to approaches based on the state-of-the-art No-U-Turn Sampler, our approach delivers substantial speedups up to an order of magnitude, while maintaining or improving predictive performance and uncertainty quantification across diverse tasks and data modalities. The suggested Microcanonical Langevin Ensembles and modifications to MCLMC additionally enhance the method's predictability in resource requirements, facilitating easier parallelization. All in all, the proposed method offers a promising direction for practical, scalable inference for BNNs.

cs.LG

An Analytic Hybrid Halo + Perturbation Theory Model for Small-scale Correlators: Baryons, Halos, and Galaxies

We update Halo Zeldovich Perturbation Theory (HZPT), an analytic model for the two-point statistics of dark matter, to describe halo and galaxy clustering, and galaxy-matter cross-correlation on nonlinear scales. The model correcting Zeldovich has an analytic Fourier transform, and therefore is valid in both configuration space and Fourier space. The model is accurate at the $2\%$-level or less for $P_{mm}$ (k < 1 h/Mpc), $P_{hm}$ (k < 1 h/Mpc), $P_{hh}$ (k < 2 h/Mpc), $P_{gm}$ (k < 1 h/Mpc), $P_{gg}$ (k < 1 h/Mpc), $ξ_{mm}$ (r > 1 Mpc/h), $ξ_{hm}$ (r > 2 Mpc/h), $ξ_{hh}$ (r > 2 Mpc/h), $ξ_{gm}$ (r > 1 Mpc/h), $ξ_{gg}$ (r > 2 Mpc/h), for LRG-like mock galaxies. We show that the HZPT model for matter correlators can account for the effects of a wide range of baryonic feedback models and provide extended dark matter models which are of $1\% ~(3\%)$ accuracy for k < 10 (8) h/Mpc. We explicitly model the non-perturbative features of halo exclusion for the halo-halo and galaxy-galaxy correlators, as well as the presence of satellites for galaxy-matter and galaxy-galaxy correlation functions. We perform density estimation using N-body simulations and a wide range of HOD galaxy mocks to obtain correlations of model parameters with the cosmological parameters $Ω_{m}$ and $σ_{8}$. HZPT can provide a fast, interpretable, and analytic model for combined-probe analyses of redshift surveys using scales well into the non-linear regime.

astro-ph.CO

Learning to Concentrate: Multi-tracer Forecasts on Local Primordial Non-Gaussianity with Machine-Learned Bias

Local primordial non-Gaussianity (LPNG) is predicted by many non-minimal models of inflation, and creates a scale-dependent contribution to the power spectrum of large-scale structure (LSS) tracers, whose amplitude is characterized by $b_ϕ$. Knowledge of $b_ϕ$ for the observed tracer population is therefore crucial for learning about inflation from LSS. Recently, it has been shown that the relationship between linear bias $b_1$ and $b_ϕ$ for simulated halos exhibits significant secondary dependence on halo concentration. We leverage this fact to forecast multi-tracer constraints on $f_{NL}^{\mathrm{loc}}$. We train a machine learning model on observable properties of simulated Illustris-TNG galaxies to predict $b_ϕ$ for samples constructed to approximate DESI emission line galaxies (ELGs) and luminous red galaxies (LRGs). We find $σ(f_{NL}^{\mathrm{loc}}) = 2.3$, and $σ(f_{NL}^{\mathrm{loc}}) = 3.7$, respectively. These forecasted errors are roughly factors of 3, and 35\% improvements over the single-tracer case for each sample, respectively. When considering both ELGs and LRGs in their overlap region, we forecast $σ(f_{NL}^{\mathrm{loc}}) = 1.5$ is attainable with our learned model, more than a factor of 3 improvement over the single-tracer case, while the ideal split by $b_ϕ$ could reach $σ(f_{NL}^{\mathrm{loc}}) <1$. We also perform multi-tracer forecasts for upcoming spectroscopic surveys targeting LPNG (MegaMapper, SPHEREx) and show that splitting tracer samples by $b_ϕ$ can lead to an order-of-magnitude reduction in projected $σ(f_{NL}^{\mathrm{loc}})$ for these surveys.

astro-ph.CO

A comparative study of cosmological constraints from weak lensing using Convolutional Neural Networks

Weak Lensing (WL) surveys are reaching unprecedented depths, enabling the investigation of very small angular scales. At these scales, nonlinear gravitational effects lead to higher-order correlations making the matter distribution highly non-Gaussian. Extracting this information using traditional statistics has proven difficult, and Machine Learning based summary statistics have emerged as a powerful alternative. We explore the capabilities of a discriminative, Convolutional Neural Networks (CNN) based approach, focusing on parameter constraints in the ($Ω_m$, $σ_8$) cosmological parameter space. Leveraging novel training loss functions and network representations on WL mock datasets without baryons, we show that our models achieve $\sim 5$ times stronger constraints than the power spectrum, $\sim 3$ stronger constraints than peak counts, and $\sim 2$ stronger constraints than previous CNN-learned summary statistics and scattering transforms, for noise levels relevant to Rubin or Euclid. For WL convergence maps with baryonic physics, our models achieve $\sim 2.3$ times stronger constraining power than the power spectrum at these noise levels, also outperforming previous summary statistics. To further explore the possibilities of CNNs for this task, we also discuss transfer learning where we adapt pre-trained models, trained on different tasks or datasets, for cosmological inference, finding that these do not improve the performance.

astro-ph.CO

Multiscale Flow for Robust and Optimal Cosmological Analysis

We propose Multiscale Flow, a generative Normalizing Flow that creates samples and models the field-level likelihood of two-dimensional cosmological data such as weak lensing. Multiscale Flow uses hierarchical decomposition of cosmological fields via a wavelet basis, and then models different wavelet components separately as Normalizing Flows. The log-likelihood of the original cosmological field can be recovered by summing over the log-likelihood of each wavelet term. This decomposition allows us to separate the information from different scales and identify distribution shifts in the data such as unknown scale-dependent systematics. The resulting likelihood analysis can not only identify these types of systematics, but can also be made optimal, in the sense that the Multiscale Flow can learn the full likelihood at the field without any dimensionality reduction. We apply Multiscale Flow to weak lensing mock datasets for cosmological inference, and show that it significantly outperforms traditional summary statistics such as power spectrum and peak counts, as well as novel Machine Learning based summary statistics such as scattering transform and convolutional neural networks. We further show that Multiscale Flow is able to identify distribution shifts not in the training data such as baryonic effects. Finally, we demonstrate that Multiscale Flow can be used to generate realistic samples of weak lensing data.

astro-ph.CO

A field-level emulator for modeling baryonic effects across hydrodynamic simulations

We develop a new and simple method to model baryonic effects at the field level relevant for weak lensing analyses. We analyze thousands of state-of-the-art hydrodynamic simulations from the CAMELS project, each with different cosmology and strength of feedback, and we find that the cross-correlation coefficient between full hydrodynamic and N-body simulations is very close to 1 down to $k\sim10~h{\rm Mpc}^{-1}$. This suggests that modeling baryonic effects at the field level down to these scales only requires N-body simulations plus a correction to the mode's amplitude given by: $\sqrt{P_{\rm hydro}(k)/P_{\rm nbody}(k)}$. In this paper, we build an emulator for this quantity, using Gaussian processes, that is flexible enough to reproduce results from thousands of hydrodynamic simulations that have different cosmologies, astrophysics, subgrid physics, volumes, resolutions, and at different redshifts. Our emulator is accurate at the percent level and exhibits a range of validation superior to previous studies. This method and our emulator enable field-level simulation-based inference analyses and accounting for baryonic effects in weak lensing analyses.

astro-ph.CO

Deterministic Langevin Unconstrained Optimization with Normalizing Flows

We introduce a global, gradient-free surrogate optimization strategy for expensive black-box functions inspired by the Fokker-Planck and Langevin equations. These can be written as an optimization problem where the objective is the target function to maximize minus the logarithm of the current density of evaluated samples. This objective balances exploitation of the target objective with exploration of low-density regions. The method, Deterministic Langevin Optimization (DLO), relies on a Normalizing Flow density estimate to perform active learning and select proposal points for evaluation. This strategy differs qualitatively from the widely-used acquisition functions employed by Bayesian Optimization methods, and can accommodate a range of surrogate choices. We demonstrate superior or competitive progress toward objective optima on standard synthetic test functions, as well as on non-convex and multi-modal posteriors of moderate dimension. On real-world objectives, such as scientific and neural network hyperparameter optimization, DLO is competitive with state-of-the-art baselines.

cs.LG

Field-Level Inference with Microcanonical Langevin Monte Carlo

Field-level inference provides a means to optimally extract information from upcoming cosmological surveys, but requires efficient sampling of a high-dimensional parameter space. This work applies Microcanonical Langevin Monte Carlo (MCLMC) to sample the initial conditions of the Universe, as well as the cosmological parameters $σ_8$ and $Ω_m$, from simulations of cosmic structure. MCLMC is shown to be over an order of magnitude more efficient than traditional Hamiltonian Monte Carlo (HMC) for a $\sim 2.6 \times 10^5$ dimensional problem. Moreover, the efficiency of MCLMC compared to HMC greatly increases as the dimensionality increases, suggesting gains of many orders of magnitude for the dimensionalities required by upcoming cosmological surveys.

astro-ph.CO

RSD measurements from BOSS galaxy power spectrum using the halo perturbation theory model

We present growth of structure constraints from the cosmological analysis of the power spectrum multipoles of SDSS-III BOSS DR12 galaxies. We use the galaxy power spectrum model of Hand et al. (2017), which decomposes the galaxies into halo mass bins, each of which is modeled separately using the relations between halo biases and halo mass. The model combines Eulerian perturbation theory and halo model calibrated on $N$-body simulations to model the halo clustering. In this work, we also generate the covariance matrix by combining the analytic disconnected part with the empirical connected part: we smooth the connected component by selecting a few principal components and show that it achieves good agreement with the mock covariance. Our analysis differs from recent analyses in that we constrain a single parameter $fσ_8$ fixing everything else to Planck+BAO prior, thereby reducing the effects of prior volume and mismodeling. We find tight constraints on $fσ_8$: $fσ_8(z_{\mathrm{eff}}=0.38)=0.489 \pm 0.038$ and $fσ_8(z_{\mathrm{eff}}=0.61)=0.455 \pm 0.028$ at $k_{\mathrm{max}} = 0.2\ h$Mpc$^{-1}$, with an overall amplitude error of 5%, and in good agreement (within 0.3 sigma) of Planck amplitude. We discuss the sensitivity of cosmological parameter estimation to the choice of scale cuts, covariance matrix, and the inclusion of hexadecapole $P_4(k)$. We show that with $k_{\mathrm{max}} = 0.4\ h$Mpc$^{-1}$ the constraints improve considerably to an overall 3.2% amplitude error, but there is some evidence of model misspecification on MultiDark-PATCHY mocks. Choosing $k_{\mathrm{max}}$ consistently and reliably remains the main challenge of RSD analysis methods.

astro-ph.CO

Deterministic Langevin Monte Carlo with Normalizing Flows for Bayesian Inference

We propose a general purpose Bayesian inference algorithm for expensive likelihoods, replacing the stochastic term in the Langevin equation with a deterministic density gradient term. The particle density is evaluated from the current particle positions using a Normalizing Flow (NF), which is differentiable and has good generalization properties in high dimensions. We take advantage of NF preconditioning and NF based Metropolis-Hastings updates for a faster convergence. We show on various examples that the method is competitive against state of the art sampling methods.

stat.ML