SearcharxivSearch

arXiv subjects

Elena Sellentin

Publications and source records attributed to Elena Sellentin.

At least 19 recordsLinked to original sources

Cosmological constraining power of the redshifts, heights, and angular clustering of weak gravitational lensing peaks

Weak gravitational lensing (WL) peaks probe non-Gaussian information of the large-scale distribution of matter that is not captured by two-point lensing statistics. We study the cosmological potential of the height distribution, redshift distribution, and angular clustering of high-valued WL peaks using a Bayesian inference approach that mimics a $\textit{Euclid}$ analysis. We use a forthcoming dark-matter-only hypercube, which varies cosmology in a ten-dimensional space, including evolving dark energy, neutrino mass, decaying dark matter, and the running of the scalar spectral index. We find the individual WL peak statistics to be complementary, as the redshift distribution best constrains the matter density, $\Omega_\mathrm{m}$, and the dark energy equation-of-state parameters, $w_0$, and $w_a$; the height distribution and angular clustering are most sensitive to the amplitude of the primordial power spectrum, $\ln(10^{10}A_\mathrm{s})$; while combining the three statistics allows us to also probe the baryon density, $\Omega_\mathrm{b}$, and the Hubble parameter, $h$. A comparison to the shear two-point correlation function demonstrates that WL peaks alone outperform the commonly used statistics, while even tighter constraints are obtained when combining all. Considering the 2-dimensional $\Omega_\mathrm{m}-\ln(10^{10}A_\mathrm{s})$ and $w_0-w_a$ parameter planes, we find that the redshift distribution outperforms the angular clustering and peak-height distributions. We study the impact of the smoothing scale and find that, typically, the smallest scales yield the best results and the figure of merit improves by a factor $\approx2$ when combining multiple scales.

astro-ph.CO

On combining estimated and analytic covariance matrices

The statistical analysis of cosmological data often assumes a Gaussian sampling distribution and relies on covariance matrices estimated from simulations. In this setting, the likelihood function of the data is not Gaussian but is instead a multivariate Student-t distribution, arising from marginalisation over an inverse-Wishart distribution for the true covariance matrix. This framework, introduced by Sellentin & Heavens (2016) and extended by Percival et al. (2022), provides a principled drop-in replacement to the Gaussian likelihood with Hartlap correction (Hartlap et al. 2007). The latter removes bias in the precision matrix; it is still widely used, despite failing to reproduce the heavy tails of the true distribution (thus yielding inaccurate probabilities, especially in the case of tensions between datasets). In practice, cosmological analyses frequently involve additional Gaussian error contributions, for example from instrumental noise, foregrounds, super-sample covariance, or emulator uncertainties. The resulting likelihood function is a convolution of the Sellentin-Heavens or Percival likelihoods with an extra Gaussian contribution, and does not have a simple expression. In this note, we derive an accurate approximation for the combined likelihood function, another multivariate Student-t distribution which inherits the heavy tails. The parameters of the Student-t distribution are determined by matching the covariance and multivariate kurtosis to those of the true distribution. We also include a slightly more expensive but fast sampling algorithm, based on the mixture representation of the Student-t distribution, which avoids the approximation altogether, but is not the drop-in replacement for the normal Gaussian or Hartlap likelihood function that the Student-t approximation in this paper provides. (Abridged)

astro-ph.CO

Almanac: HMC sampling with bounded velocity

In Hamiltonian Monte Carlo sampling, the shape of the potential and the choice of the momentum distribution jointly give rise to the Hamiltonian dynamics of the sampler. An efficient sampler propagates quickly in all regions of the parameter space, so that the chain has a low autocorrelation length and the sampler has a high acceptance rate, with the goal of optimising the number of near-independent samples for given computational cost. Standard Gaussian momentum distributions allow arbitrarily large velocities, which can lead to inefficient exploration in posteriors with ridges or funnel-like geometries. We investigate alternative momentum distributions based on relativistic and Student's t kinetic energies, which naturally limit particle velocities and may improve robustness. Using Almanac, a sampler for cosmological posterior distributions of sky maps and power spectra on the sphere, we test these alternatives in both low- and high-dimensional settings. We find that the choice of parameterization and momentum distribution can improve convergence and effective sample rate, though the achievable gains are generally modest and strongly problem-dependent, reaching up to an order of magnitude in favorable cases. Among the momentum distributions that we tested, those with moderately heavy tails achieved the best balance between efficiency and stability. These results highlight the importance of sampler design and encourage future work on adaptive and self-tuning strategies for kinetic energy parameter optimization in high-dimensional settings.

astro-ph.CO

FLAMINGO: Baryonic effects on the weak lensing scattering transform

The scattering transform is a wavelet-based statistic capable of capturing non-Gaussian features in weak lensing (WL) convergence maps and has been proven to tighten cosmological parameter constraints by accessing information beyond two-point functions. However, its application in cosmological inference requires a clear understanding of its sensitivity to astrophysical systematics, the most significant of which are baryonic effects. These processes substantially modify the matter distribution on small to intermediate scales ($k\gtrsim 0.1\,h\,\mathrm{Mpc}^{-1}$), leaving scale-dependent imprints on the WL convergence field. We systematically examine the impact of baryonic feedback on scattering coefficients using full-sky WL convergence maps with Stage IV survey characteristics, generated from the FLAMINGO simulation suite. These simulations include a broad range of feedback models, calibrated to match the observed cluster gas fraction and galaxy stellar mass function, including systematically shifted variations, and incorporating either thermal or jet-mode AGN feedback. We characterise baryonic effects using a baryonic transfer function defined as the ratio of hydrodynamical to dark-matter-only scattering coefficients. While the coefficients themselves are sensitive to both cosmology and feedback, the transfer function remains largely insensitive to cosmology and shows a strong response to feedback, with suppression reaching up to $10\%$ on scales of $k\gtrsim 0.1\,h\,\mathrm{Mpc}^{-1}$. We also demonstrate that shape noise significantly diminishes the sensitivity of the scattering coefficients to baryonic effects, reducing the suppression from $\sim 2 - 10 \;\%$ to $\sim 1\;\%$, even with 1.5 arcmin Gaussian smoothing. This highlights the need for noise mitigation strategies and high-resolution data in future WL surveys.

astro-ph.CO

Stage IV baryonic feedback correction for non-Gaussianity inference

Non-Gaussian statistics of the projected weak lensing field are powerful estimators that can outperform the constraining power of the two-point functions in inferring cosmological parameters. This is because these estimators extract the non-Gaussian information contained in the small scales. However, fully leveraging the statistical precision of such estimators is hampered by theoretical uncertainties, such as those arising from baryonic physics. Moreover, as non-Gaussian estimators mix different scales, there exists no natural cut-off scale below which baryonic feedback can be completely removed. We therefore present a Bayesian solution for accounting for baryonic feedback uncertainty in weak lensing non-Gaussianity inference. Our solution implements Bayesian model averaging (BMA), a statistical framework that accounts for model uncertainty and combines the strengths of different models to produce more robust and reliable parameter inferences. We demonstrate the effectiveness of this approach in a Stage IV convergence peak counts analysis, including three baryonic feedback models. We find that the resulting BMA posterior distribution safeguards parameter inference against biases due to baryonic feedback, and therefore provides a robust framework for obtaining accurate cosmological constraints at Stage IV precision under model uncertainty scenarios.

astro-ph.CO

The distribution of Bayes' ratio

The ratio of Bayesian evidences is a popular tool in cosmology to compare different models. There are however several issues with this method: Bayes' ratio depends on the prior even in the limit of non-informative priors, and Jeffrey's scale, used to assess the test, is arbitrary. Moreover, the standard use of Bayes' ratio is often criticized for being unable to reject models. In this paper, we address these shortcoming by promoting evidences and evidence ratios to frequentist statistics and deriving their sampling distributions. By comparing the evidence ratios to their sampling distributions, poor fitting models can now be rejected. Our method additionally does not depend on the prior in the limit of very weak priors, thereby safeguarding the experimenter against premature rejection of a theory with a uninformative prior, and replaces the arbitrary Jeffrey's scale by probability thresholds for rejection. We provide analytical solutions for some simplified cases (Gaussian data, linear parameters, and nested models), and we apply the method to cosmological supernovae Ia data. We dub our method the FB method, for Frequentist-Bayesian.

astro-ph.CO

A hybrid approach for solving the gravitational N-body problem with Artificial Neural Networks

Simulating the evolution of the gravitational N-body problem becomes extremely computationally expensive as N increases since the problem complexity scales quadratically with the number of bodies. We study the use of Artificial Neural Networks (ANNs) to replace expensive parts of the integration of planetary systems. Neural networks that include physical knowledge have grown in popularity in the last few years, although few attempts have been made to use them to speed up the simulation of the motion of celestial bodies. We study the advantages and limitations of using Hamiltonian Neural Networks to replace computationally expensive parts of the numerical simulation. We compare the results of the numerical integration of a planetary system with asteroids with those obtained by a Hamiltonian Neural Network and a conventional Deep Neural Network, with special attention to understanding the challenges of this problem. Due to the non-linear nature of the gravitational equations of motion, errors in the integration propagate. To increase the robustness of a method that uses neural networks, we propose a hybrid integrator that evaluates the prediction of the network and replaces it with the numerical solution if considered inaccurate. Hamiltonian Neural Networks can make predictions that resemble the behavior of symplectic integrators but are challenging to train and in our case fail when the inputs differ ~7 orders of magnitude. In contrast, Deep Neural Networks are easy to train but fail to conserve energy, leading to fast divergence from the reference solution. The hybrid integrator designed to include the neural networks increases the reliability of the method and prevents large energy errors without increasing the computing cost significantly. For this problem, the use of neural networks results in faster simulations when the number of asteroids is >70.

astro-ph.EP

Extreme data compression for Bayesian model comparison

We develop extreme data compression for use in Bayesian model comparison via the MOPED algorithm, as well as more general score compression. We find that Bayes factors from data compressed with the MOPED algorithm are identical to those from their uncompressed datasets when the models are linear and the errors Gaussian. In other nonlinear cases, whether nested or not, we find negligible differences in the Bayes factors, and show this explicitly for the Pantheon-SH0ES supernova dataset. We also investigate the sampling properties of the Bayesian Evidence as a frequentist statistic, and find that extreme data compression reduces the sampling variance of the Evidence, but has no impact on the sampling distribution of Bayes factors. Since model comparison can be a very computationally-intensive task, MOPED extreme data compression may present significant advantages in computational time.

astro-ph.IM

Extremely expensive likelihoods: A variational-Bayes solution for precision cosmology

We present a variational-Bayes solution to compute non-Gaussian posteriors from extremely expensive likelihoods. Our approach is an alternative for parameter inference when MCMC sampling is numerically prohibitive or conceptually unfeasible. For example, when either the likelihood or the theoretical model cannot be evaluated at arbitrary parameter values, but only previously selected values, then traditional MCMC sampling is impossible, whereas our variational-Bayes solution still succeeds in estimating the full posterior. In cosmology, this occurs e.g. when the parametric model is based on costly simulations that were run for previously selected input parameters. We demonstrate the applicability of our posterior construction on the KiDS-450 weak lensing analysis, where we reconstruct the original KiDS MCMC posterior at 0.6% of its former numerical posterior evaluations. The reduction in numerical cost implies that systematic effects which formerly exhausted the numerical budget could now be included.

astro-ph.CO

Identifying the most constraining ice observations to infer molecular binding energies

In order to understand grain-surface chemistry, one must have a good understanding of the reaction rate parameters. For diffusion-based reactions, these parameters are binding energies of the reacting species. However, attempts to estimate these values from grain-surface abundances using Bayesian inference are inhibited by a lack of enough sufficiently constraining data. In this work, we use the Massive Optimised Parameter Estimation and Data (MOPED) compression algorithm to determine which species should be prioritised for future ice observations to better constrain molecular binding energies. Using the results from this algorithm, we make recommendations for which species future observations should focus on.

astro-ph.GA

Bayesian error propagation for neural-net based parameter inference

Neural nets have become popular to accelerate parameter inferences, especially for the upcoming generation of galaxy surveys in cosmology. As neural nets are approximative by nature, a recurrent question has been how to propagate the neural net's approximation error, in order to avoid biases in the parameter inference. We present a Bayesian solution to propagating a neural net's approximation error and thereby debiasing parameter inference. We exploit that a neural net reports its approximation errors during the validation phase. We capture the thus reported approximation errors via the highest-order summary statistics, allowing us to eliminate the neural net's bias during inference, and propagating its uncertainties. We demonstrate that our method is quickly implemented and successfully infers parameters even for strongly biased neural nets. In summary, our method provides the missing element to judge the accuracy of a posterior if it cannot be computed based on an infinitely accurately theory code.

astro-ph.IM

Impact of bar resonances in the velocity-space distribution of the solar neighbourhood stars in a self-consistent $N$-body Galactic disc simulation

The velocity-space distribution of the solar neighbourhood stars shows complex substructures. Most of the previous studies use static potentials to investigate their origins. Instead we use a self-consistent $N$-body model of the Milky Way, whose potential is asymmetric and evolves with time. In this paper, we quantitatively evaluate the similarities of the velocity-space distributions in the $N$-body model and that of the solar neighbourhood, using Kullback-Leibler divergence (KLD). The KLD analysis shows the time evolution and spatial variation of the velocity-space distribution. The KLD fluctuates with time, which indicates the velocity-space distribution at a fixed position is not always similar to that of the solar neighbourhood. Some positions show velocity-space distributions with small KLDs (high similarities) more frequently than others. One of them locates at $(R,ϕ)=(8.2\;\mathrm{kpc}, 30^{\circ})$, where $R$ and $ϕ$ are the distance from the galactic centre and the angle with respect to the bar's major axis, respectively. The detection frequency is higher in the inter-arm regions than in the arm regions. In the velocity maps with small KLDs, we identify the velocity-space substructures, which consist of particles trapped in bar resonances. The bar resonances have significant impact on the stellar velocity-space distribution even though the galactic potential is not static.

astro-ph.GA

Ultra-large-scale approximations and galaxy clustering: debiasing constraints on cosmological parameters

Upcoming galaxy surveys will allow us to probe the growth of the cosmic large-scale structure with improved sensitivity compared to current missions, and will also map larger areas of the sky. This means that in addition to the increased precision in observations, future surveys will also access the ultra-large-scale regime, where commonly neglected effects such as lensing, redshift-space distortions and relativistic corrections become important for calculating correlation functions of galaxy positions. At the same time, several approximations usually made in these calculations, such as the Limber approximation, break down at those scales. The need to abandon these approximations and simplifying assumptions at large scales creates severe issues for parameter estimation methods. On the one hand, exact calculations of theoretical angular power spectra become computationally expensive, and the need to perform them thousands of times to reconstruct posterior probability distributions for cosmological parameters makes the approach unfeasible. On the other hand, neglecting relativistic effects and relying on approximations may significantly bias the estimates of cosmological parameters. In this work, we quantify this bias and investigate how an incomplete modelling of various effects on ultra-large scales could lead to false detections of new physics beyond the standard $Λ$CDM model. Furthermore, we propose a simple debiasing method that allows us to recover true cosmologies without running the full parameter estimation pipeline with exact theoretical calculations. This method can therefore provide a fast way of obtaining accurate values of cosmological parameters and estimates of exact posterior probability distributions from ultra-large-scale observations.

astro-ph.CO

MCMC generation of cosmological fields far beyond Gaussianity

Structure formation in our Universe creates non-Gaussian random fields that will soon be observed over almost the entire sky by the Euclid satellite, the Vera-Rubin observatory, and the Square Kilometre Array. An unsolved problem is how to analyze best such non-Gaussian fields, e.g. to infer the physical laws that created them. This problem could be solved if a parametric non-Gaussian sampling distribution for such fields were known, as this distribution could serve as likelihood during inference. We therefore create a sampling distribution for non-Gaussian random fields. Our approach is capable of handling strong non-Gaussianity, while perturbative approaches such as the Edgeworth expansion cannot. To imitate cosmological structure formation, we enforce our fields to be (i) statistically isotropic, (ii) statistically homogeneous, and (iii) statistically independent at large distances. We generate such fields via a Monte Carlo Markov Chain technique and find that even strong non-Gaussianity is not necessarily visible to the human eye. We also find that sampled marginals for pixel pairs have an almost generic Gauss-like appearance, even if the joint distribution of all pixels is markedly non-Gaussian. This apparent Gaussianity is a consequence of the high dimensionality of random fields. We conclude that vast amounts of non-Gaussian information can be hidden in random fields that appear nearly Gaussian in simple tests, and that it would be short-sighted not to try and extract it.

astro-ph.CO

Matching Bayesian and frequentist coverage probabilities when using an approximate data covariance matrix

Observational astrophysics consists of making inferences about the Universe by comparing data and models. The credible intervals placed on model parameters are often as important as the maximum a posteriori probability values, as the intervals indicate concordance or discordance between models and with measurements from other data. Intermediate statistics (e.g. the power spectrum) are usually measured and inferences made by fitting models to these rather than the raw data, assuming that the likelihood for these statistics has multivariate Gaussian form. The covariance matrix used to calculate the likelihood is often estimated from simulations, such that it is itself a random variable. This is a standard problem in Bayesian statistics, which requires a prior to be placed on the true model parameters and covariance matrix, influencing the joint posterior distribution. As an alternative to the commonly-used Independence-Jeffreys prior, we introduce a prior that leads to a posterior that has approximately frequentist matching coverage. This is achieved by matching the covariance of the posterior to that of the distribution of true values of the parameters around the maximum likelihood values in repeated trials, under certain assumptions. Using this prior, credible intervals derived from a Bayesian analysis can be interpreted approximately as confidence intervals, containing the truth a certain proportion of the time for repeated trials. Linking frequentist and Bayesian approaches that have previously appeared in the astronomical literature, this offers a consistent and conservative approach for credible intervals quoted on model parameters for problems where the covariance matrix is itself an estimate.

astro-ph.IM

Galactic potential constraints from clustering in action space of combined stellar stream data

Stream stars removed by tides from their progenitor satellite galaxy or globular cluster act as a group of test particles on neighboring orbits, probing the gravitational field of the Milky Way. While constraints from individual streams have been shown to be susceptible to biases, combining several streams from orbits with various distances reduces these biases. We fit a common gravitational potential to multiple stellar streams simultaneously by maximizing the clustering of the stream stars in action space. We apply this technique to members of the GD-1, Pal 5, Orphan and Helmi streams, exploiting both the individual and combined data sets. We describe the Galactic potential with a Stäckel model, and vary up to five parameters simultaneously. We find that we can only constrain the enclosed mass, and that the strongest constraints come from the GD-1, Pal 5 and Orphan streams whose combined data set yields $M(< 20\ \mathrm{kpc}) = 2.96^{+0.25}_{-0.26} \times 10^{11} \ M_{\odot}$. When including the Helmi stream in the data set, the mass uncertainty increases to $M(< 20\ \mathrm{kpc}) = 3.12^{+3.21}_{-0.46} \times 10^{11} \ M_{\odot}$.

astro-ph.GA

The impact of signal-to-noise, redshift, and angular range on the bias of weak lensing 2-point functions

Weak lensing data follow a naturally skewed distribution, implying the data vector most likely yielded from a survey will systematically fall below its mean. Although this effect is qualitatively known from CMB-analyses, correctly accounting for it in weak lensing is challenging, as a direct transfer of the CMB results is quantitatively incorrect. While a previous study (Sellentin et al. 2018) focused on the magnitude of this bias, we here focus on the frequency of this bias, its scaling with redshift, and its impact on the signal-to-noise of a survey. Filtering weak lensing data with COSEBIs, we show that weak lensing likelihoods are skewed up until $\ell \approx 100$, whereas CMB-likelihoods Gaussianize already at $\ell \approx 20$. While COSEBI-compressed data on KiDS- and DES-like redshift- and angular ranges follow Gaussian distributions, we detect skewness at 6$σ$ significance for half of a Euclid- or LSST-like data set, caused by the wider coverage and deeper reach of these surveys. Computing the signal-to-noise ratio per data point, we show that precisely the data points of highest signal-to-noise are the most biased. Over all redshifts, this bias affects at least 10% of a survey's total signal-to-noise, at high redshifts up to 25%. The bias is accordingly expected to impact parameter inference. The bias can be handled by developing non-Gaussian likelihoods. Otherwise, it could be reduced by removing the data points of highest signal-to-noise.

astro-ph.CO

Trimodal structure of Hercules stream explained by originating from bar resonances

Gaia Data Release 2 revealed detailed structures of nearby stars in phase space. These include the Hercules stream, whose origin is still debated. Most of the previous numerical studies conjectured that the observed structures originate from orbits in resonance with the bar, based on static potential models for the Milky Way. We, in contrast, approach the problem via a self-consistent, dynamic, and morphologically well-resolved model, namely a full $N$-body simulation of the Milky Way. Our simulation comprises about 5.1 billion particles in the galactic stellar bulge, bar, disk, and dark-matter halo and is evolved to 10 Gyr. Our model's disk component is composed of 200 million particles, and its simulation snapshots are stored every 10 Myr, enabling us to resolve and classify resonant orbits of representative samples of stars. After choosing the Sun's position in the simulation, we compare the distribution of stars in its neighborhood with Gaia's astrometric data, thereby establishing the role of identified resonantly trapped stars in the formation of Hercules-like structures. From our orbital spectral-analysis we identify multiple, especially higher order resonances. Our results suggest that the Hercules stream is dominated by the 4:1 and 5:1 outer Lindblad and corotation resonances. In total, this yields a trimodal structure of the Hercules stream. From the relation between resonances and ridges in phase space, our model favored a slow pattern speed of the Milky-Way bar (40--45 $\mathrm{km \; s^{-1} \; kpc^{-1}}$).

astro-ph.GA