SearcharxivSearch

arXiv subjects

Niall Jeffrey

Publications and source records attributed to Niall Jeffrey.

At least 19 recordsLinked to original sources

Field-level weak lensing cosmology with $60$ simulations using multifidelity simulation-based inference

We perform a realistic KiDS-Legacy mock analysis with field-level neural compression and simulation-based inference using just 60 $N$-body simulations. The weak lensing shear field encodes substantially more cosmological information than standard two-point summary statistics such as the power spectrum. Field-level inference can fully exploit this information, but physical realism at the field-level requires very high-fidelity simulations. This poses a major challenge for simulation-based inference (SBI): accurate empirical density modelling and deep-learning-based neural compression require tens of thousands of training samples, but achieving physical realism at the field level makes each simulation extremely costly. We demonstrate that multifidelity SBI can alleviate this tension by substantially reducing the number of high-fidelity simulations needed for accurate cosmological inference. We pre-train neural inference models on realistic KiDS-Legacy-like shear mocks using fast log-normal \texttt{GLASS} simulations and fine-tune them on a small set of high-fidelity $N$-body simulations. We show that $60$ high-fidelity simulations are sufficient to obtain informative and well-calibrated cosmological posteriors, enabling at least an order-of-magnitude reduction in simulation cost for accurate field-level inference in a realistic setting.

astro-ph.CO

The Degeneracy Distillery

When two or more parameters or labels produce similar data, they are degenerate, or hard to distinguish. Degeneracies render both label prediction and inverse problems difficult, since both machine learning algorithms and probabilistic samplers rely on the distinguishability of data and its gradients with respect to parameters. However, identifying degeneracies in physical models or real-world datasets can be elucidating about the choice of model or the underlying process that produces the data. We present the degeneracy distillery, a method that (1) detects and (2) resolves degenerate parameter combinations (a) automatically and (b) symbolically, from parameter-data (or parameter-simulation) pairs alone, through estimation and flattening of the Fisher information matrix. By exploring the information geometry of the likelihood, we characterize degeneracies as an intrinsic property of the physical model, requiring no realised data observation. We demonstrate our approach on a range of synthetic and real-world problems, discovering symbolic coordinate transformations that identify the combinations of parameters of a model which yield independent effects on the data. The resulting coordinates flatten the Fisher information in expectation globally, in contrast to posterior-based methods that flatten only at a single point, and substantially reduce the simulation budget required for downstream neural posterior estimation. In test cases we require up to $10\times$ fewer simulations for posterior estimation at matched validation calibration whilst simultaneously gaining physical insight on the system.

cs.LG

Intrinsic galaxy alignments in CAMELS simulations and the significant impact of baryon model

We present a detection of the intrinsic galaxy alignments in the CAMELS suite of hydrodynamic simulations. We find that the alignment amplitude depends significantly on cosmological and supernova feedback parameters - specifically $\Omega_m$, $\sigma_8$, $A_{\text{SN1}}$, $A_{\text{SN2}}$- while no dependence on AGN feedback is observed (due to the limited simulation volume $(25\,h^{-1}\,\text{Mpc})^3$). The dependence on $\sigma_8$ vanishes when projected correlation functions $w_{m+}$ are normalized by matter density correlations $w_{mm}$, consistent with predictions from linear alignment models. We find alignment amplitudes in quiescent galaxies to exceed those in star-forming galaxies by an order of magnitude. Moreover, examining orientation-only correlation functions from ellipticity-normalized galaxies $\tilde w_{m+}$, we confirm that alignment signals retain sensitivity to supernova feedback across full, star forming, quiescent, and ellipticity-normalized samples. Finally, we find evidence that supernova feedback impacts alignment signals differently in star-forming versus quiescent populations, suggesting that distinct alignment mechanisms operate across galaxy types. Our results offer key insights for understanding galaxy formation and alignment models for future weak gravitational lensing analyses.

astro-ph.CO

Searching for Dark Structures: A Comparison of Weak-lensing Convergence Maps and Lensing-weighted Galaxy Density Maps

We present the result of a comparison between the dark matter distribution inferred from weak gravitational lensing and the observed galaxy distribution to identify dark structures with a high dark matter-to-galaxy density ratio. To do this, we use weak-lensing convergence maps from the Dark Energy Survey Year 3 data, and construct corresponding galaxy convergence maps at $z\lesssim1.0$ , representing projected galaxy number density fluctuations weighted by lensing efficiency. The two maps show overall agreement. However, we could identify 22 regions where the dark matter density exhibits an excess compared to the galaxy density. After carefully examining the survey depths and proximity to survey boundaries, we select the seven most probable candidates for dark structures. This sample provides valuable test beds for further investigations into dark matter mapping. Moreover, our method will be very useful for future studies of dark structures as large-scale weak-lensing surveys become available, such as the Euclid mission, the Vera C. Rubin Observatory's Legacy Survey of Space and Time, and the Nancy Grace Roman Space Telescope.

astro-ph.CO

Multilevel neural simulation-based inference

Neural simulation-based inference (SBI) is a popular set of methods for Bayesian inference when models are only available in the form of a simulator. These methods are widely used in the sciences and engineering, where writing down a likelihood can be significantly more challenging than constructing a simulator. However, the performance of neural SBI can suffer when simulators are computationally expensive, thereby limiting the number of simulations that can be performed. In this paper, we propose a novel approach to neural SBI which leverages multilevel Monte Carlo techniques for settings where several simulators of varying cost and fidelity are available. We demonstrate through both theoretical analysis and extensive experiments that our method can significantly enhance the accuracy of SBI methods given a fixed computational budget.

stat.ML

Transfer learning for multifidelity simulation-based inference in cosmology

Simulation-based inference (SBI) enables cosmological parameter estimation when closed-form likelihoods or models are unavailable. However, SBI relies on machine learning for neural compression and density estimation. This requires large training datasets which are prohibitively expensive for high-quality simulations. We overcome this limitation with multifidelity transfer learning, combining less expensive, lower-fidelity simulations with a limited number of high-fidelity simulations. We demonstrate our methodology on dark matter density maps from two separate simulation suites in the hydrodynamical CAMELS Multifield Dataset. Pre-training on dark-matter-only $N$-body simulations reduces the required number of high-fidelity hydrodynamical simulations by a factor between $8$ and $15$, depending on the model complexity, posterior dimensionality, and performance metrics used. By leveraging cheaper simulations, our approach enables performant and accurate inference on high-fidelity models while substantially reducing computational costs.

astro-ph.CO

The Gravitational Lensing Imprints of DES Y3 Superstructures on the CMB: A Matched Filtering Approach

$ $Low density cosmic voids gravitationally lens the cosmic microwave background (CMB), leaving a negative imprint on the CMB convergence $\kappa$. This effect provides insight into the distribution of matter within voids, and can also be used to study the growth of structure. We measure this lensing imprint by cross-correlating the Planck CMB lensing convergence map with voids identified in the Dark Energy Survey Year 3 data set, covering approximately 4,200 deg$^2$ of the sky. We use two distinct void-finding algorithms: a 2D void-finder which operates on the projected galaxy density field in thin redshift shells, and a new code, Voxel, which operates on the full 3D map of galaxy positions. We employ an optimal matched filtering method for cross-correlation, using the MICE N-body simulation both to establish the template for the matched filter and to calibrate detection significances. Using the DES Y3 photometric luminous red galaxy sample, we measure $A_\kappa$, the amplitude of the observed lensing signal relative to the simulation template, obtaining $A_\kappa = 1.03 \pm 0.22$ ($4.6\sigma$ significance) for Voxel and $A_\kappa = 1.02 \pm 0.17$ ($5.9\sigma$ significance) for 2D voids, both consistent with $\Lambda$CDM expectations. We additionally invert the 2D void-finding process to identify superclusters in the projected density field, for which we measure $A_\kappa = 0.87 \pm 0.15$ ($5.9\sigma$ significance). The leading source of noise in our measurements is Planck noise, implying that future data from the Atacama Cosmology Telescope (ACT), South Pole Telescope (SPT) and CMB-S4 will increase sensitivity and allow for more precise measurements.

astro-ph.CO

Evidence Networks: simple losses for fast, amortized, neural Bayesian model comparison

Evidence Networks can enable Bayesian model comparison when state-of-the-art methods (e.g. nested sampling) fail and even when likelihoods or priors are intractable or unknown. Bayesian model comparison, i.e. the computation of Bayes factors or evidence ratios, can be cast as an optimization problem. Though the Bayesian interpretation of optimal classification is well-known, here we change perspective and present classes of loss functions that result in fast, amortized neural estimators that directly estimate convenient functions of the Bayes factor. This mitigates numerical inaccuracies associated with estimating individual model probabilities. We introduce the leaky parity-odd power (l-POP) transform, leading to the novel ``l-POP-Exponential'' loss function. We explore neural density estimation for data probability in different models, showing it to be less accurate and scalable than Evidence Networks. Multiple real-world and synthetic examples illustrate that Evidence Networks are explicitly independent of dimensionality of the parameter space and scale mildly with the complexity of the posterior probability density function. This simple yet powerful approach has broad implications for model inference tasks. As an application of Evidence Networks to real-world data we compute the Bayes factor for two models with gravitational lensing data of the Dark Energy Survey. We briefly discuss applications of our methods to other, related problems of model comparison and evaluation in implicit inference settings.

cs.LG

GLASS: Generator for Large Scale Structure

We present GLASS, the Generator for Large Scale Structure, a new code for the simulation of galaxy surveys for cosmology, which iteratively builds a light cone with matter, galaxies, and weak gravitational lensing signals as a sequence of nested shells. This allows us to create deep and realistic simulations of galaxy surveys at high angular resolution on standard computer hardware and with low resource consumption. GLASS also introduces a new technique to generate transformations of Gaussian random fields (including lognormal) to essentially arbitrary precision, an iterative line-of-sight integration over matter shells to obtain weak lensing fields, and flexible modelling of the galaxies sector. We demonstrate that GLASS readily produces simulated data sets with per cent-level accurate two-point statistics of galaxy clustering and weak lensing, thus enabling simulation-based validation and inference that is limited only by our current knowledge of the input matter and galaxy properties.

astro-ph.CO

Rediscovering orbital mechanics with machine learning

We present an approach for using machine learning to automatically discover the governing equations and hidden properties of real physical systems from observations. We train a "graph neural network" to simulate the dynamics of our solar system's Sun, planets, and large moons from 30 years of trajectory data. We then use symbolic regression to discover an analytical expression for the force law implicitly learned by the neural network, which our results showed is equivalent to Newton's law of gravitation. The key assumptions that were required were translational and rotational equivariance, and Newton's second and third laws of motion. Our approach correctly discovered the form of the symbolic force law. Furthermore, our approach did not require any assumptions about the masses of planets and moons or physical constants. They, too, were accurately inferred through our methods. Though, of course, the classical law of gravitation has been known since Isaac Newton, our result serves as a validation that our method can discover unknown laws and hidden properties from observed data. More broadly this work represents a key step toward realizing the potential of machine learning for accelerating scientific discovery.

astro-ph.EP

Probabilistic Mass Mapping with Neural Score Estimation

Weak lensing mass-mapping is a useful tool to access the full distribution of dark matter on the sky, but because of intrinsic galaxy ellipticies and finite fields/missing data, the recovery of dark matter maps constitutes a challenging ill-posed inverse problem. We introduce a novel methodology allowing for efficient sampling of the high-dimensional Bayesian posterior of the weak lensing mass-mapping problem, and relying on simulations for defining a fully non-Gaussian prior. We aim to demonstrate the accuracy of the method on simulations, and then proceed to applying it to the mass reconstruction of the HST/ACS COSMOS field. The proposed methodology combines elements of Bayesian statistics, analytic theory, and a recent class of Deep Generative Models based on Neural Score Matching. This approach allows us to do the following: 1) Make full use of analytic cosmological theory to constrain the 2pt statistics of the solution. 2) Learn from cosmological simulations any differences between this analytic prior and full simulations. 3) Obtain samples from the full Bayesian posterior of the problem for robust Uncertainty Quantification. We demonstrate the method on the $κ$TNG simulations and find that the posterior mean significantly outperfoms previous methods (Kaiser-Squires, Wiener filter, Sparsity priors) both on root-mean-square error and in terms of the Pearson correlation. We further illustrate the interpretability of the recovered posterior by establishing a close correlation between posterior convergence values and SNR of clusters artificially introduced into a field. Finally, we apply the method to the reconstruction of the HST/ACS COSMOS field and yield the highest quality convergence map of this field to date.

astro-ph.CO

Single frequency CMB B-mode inference with realistic foregrounds from a single training image

With a single training image and using wavelet phase harmonic augmentation, we present polarized Cosmic Microwave Background (CMB) foreground marginalization in a high-dimensional likelihood-free (Bayesian) framework. We demonstrate robust foreground removal using only a single frequency of simulated data for a BICEP-like sky patch. Using Moment Networks we estimate the pixel-level posterior probability for the underlying {E,B} signal and validate the statistical model with a quantile-type test using the estimated marginal posterior moments. The Moment Networks use a hierarchy of U-Net convolutional neural networks. This work validates such an approach in the most difficult limiting case: pixel-level, noise-free, highly non-Gaussian dust foregrounds with a single training image at a single frequency. For a real CMB experiment, a small number of representative sky patches would provide the training data required for full cosmological inference. These results enable robust likelihood-free, simulation-based parameter and model inference for primordial B-mode detection using observed CMB polarization data.

astro-ph.CO

A new approach for the statistical denoising of Planck interstellar dust polarization data

Dust emission is the main foreground for cosmic microwave background (CMB) polarization. Its statistical characterization must be derived from the analysis of observational data because the precision required for a reliable component separation is far greater than what is currently achievable with physical models of the turbulent magnetized interstellar medium. This letter takes a significant step toward this goal by proposing a method that retrieves non-Gaussian statistical characteristics of dust emission from noisy Planck polarization observations at 353 GHz. We devised a statistical denoising method based on wavelet phase harmonics (WPH) statistics, which characterize the coherent structures in non-Gaussian random fields and define a generative model of the data. The method was validated on mock data combining a dust map from a magnetohydrodynamic simulation and Planck noise maps. The denoised map reproduces the true power spectrum down to scales where the noise power is an order of magnitude larger than that of the signal. It remains highly correlated to the true emission and retrieves some of its non-Gaussian properties. Applied to Planck data, the method provides a new approach to building a generative model of dust polarization that will characterize the full complexity of the dust emission. We also release PyWPH, a public Python package, to perform GPU-accelerated WPH analyses on images.

astro-ph.CO

The sum of the masses of the Milky Way and M31: a likelihood-free inference approach

We use Density Estimation Likelihood-Free Inference, $Λ$ Cold Dark Matter simulations of $\sim 2M$ galaxy pairs, and data from Gaia and the Hubble Space Telescope to infer the sum of the masses of the Milky Way and Andromeda (M31) galaxies, the two main components of the Local Group. This method overcomes most of the approximations of the traditional timing argument, makes the writing of a theoretical likelihood unnecessary, and allows the non-linear modelling of observational errors that take into account correlations in the data and non-Gaussian distributions. We obtain an $M_{200}$ mass estimate $M_{\rm MW+M31} = 4.6^{+2.3}_{-1.8} \times 10^{12} M_{\odot}$ ($68 \%$ C.L.), in agreement with previous estimates both for the sum of the two masses and for the individual masses. This result is not only one of the most reliable estimates of the sum of the two masses to date, but is also an illustration of likelihood-free inference in a problem with only one parameter and only three data points.

astro-ph.GA

Probabilistic Mapping of Dark Matter by Neural Score Matching

The Dark Matter present in the Large-Scale Structure of the Universe is invisible, but its presence can be inferred through the small gravitational lensing effect it has on the images of far away galaxies. By measuring this lensing effect on a large number of galaxies it is possible to reconstruct maps of the Dark Matter distribution on the sky. This, however, represents an extremely challenging inverse problem due to missing data and noise dominated measurements. In this work, we present a novel methodology for addressing such inverse problems by combining elements of Bayesian statistics, analytic physical theory, and a recent class of Deep Generative Models based on Neural Score Matching. This approach allows to do the following: (1) make full use of analytic cosmological theory to constrain the 2pt statistics of the solution, (2) learn from cosmological simulations any differences between this analytic prior and full simulations, and (3) obtain samples from the full Bayesian posterior of the problem for robust Uncertainty Quantification. We present an application of this methodology on the first deep-learning-assisted Dark Matter map reconstruction of the Hubble Space Telescope COSMOS field.

astro-ph.CO

Likelihood-free inference with neural compression of DES SV weak lensing map statistics

In many cosmological inference problems, the likelihood (the probability of the observed data as a function of the unknown parameters) is unknown or intractable. This necessitates approximations and assumptions, which can lead to incorrect inference of cosmological parameters, including the nature of dark matter and dark energy, or create artificial model tensions. Likelihood-free inference covers a novel family of methods to rigorously estimate posterior distributions of parameters using forward modelling of mock data. We present likelihood-free cosmological parameter inference using weak lensing maps from the Dark Energy Survey (DES) SV data, using neural data compression of weak lensing map summary statistics. We explore combinations of the power spectra, peak counts, and neural compressed summaries of the lensing mass map using deep convolution neural networks. We demonstrate methods to validate the inference process, for both the data modelling and the probability density estimation steps. Likelihood-free inference provides a robust and scalable alternative for rigorous large-scale cosmological inference with galaxy survey data (for DES, Euclid and LSST). We have made our simulated lensing maps publicly available.

astro-ph.CO

Solving high-dimensional parameter inference: marginal posterior densities & Moment Networks

High-dimensional probability density estimation for inference suffers from the "curse of dimensionality". For many physical inference problems, the full posterior distribution is unwieldy and seldom used in practice. Instead, we propose direct estimation of lower-dimensional marginal distributions, bypassing high-dimensional density estimation or high-dimensional Markov chain Monte Carlo (MCMC) sampling. By evaluating the two-dimensional marginal posteriors we can unveil the full-dimensional parameter covariance structure. We additionally propose constructing a simple hierarchy of fast neural regression models, called Moment Networks, that compute increasing moments of any desired lower-dimensional marginal posterior density; these reproduce exact results from analytic posteriors and those obtained from Masked Autoregressive Flows. We demonstrate marginal posterior density estimation using high-dimensional LIGO-like gravitational wave time series and describe applications for problems of fundamental cosmology.

stat.ML

Deep learning dark matter map reconstructions from DES SV weak lensing data

We present the first reconstruction of dark matter maps from weak lensing observational data using deep learning. We train a convolution neural network (CNN) with a Unet based architecture on over $3.6\times10^5$ simulated data realizations with non-Gaussian shape noise and with cosmological parameters varying over a broad prior distribution. We interpret our newly created DES SV map as an approximation of the posterior mean $P(κ| γ)$ of the convergence given observed shear. Our DeepMass method is substantially more accurate than existing mass-mapping methods. With a validation set of 8000 simulated DES SV data realizations, compared to Wiener filtering with a fixed power spectrum, the DeepMass method improved the mean-square-error (MSE) by 11 per cent. With N-body simulated MICE mock data, we show that Wiener filtering with the optimal known power spectrum still gives a worse MSE than our generalized method with no input cosmological parameters; we show that the improvement is driven by the non-linear structures in the convergence. With higher galaxy density in future weak lensing data unveiling more non-linear scales, it is likely that deep learning will be a leading approach for mass mapping with Euclid and LSST.

astro-ph.CO