SearcharxivSearch

arXiv subjects

Kiyam Lin

Publications and source records attributed to Kiyam Lin.

7 recordsLinked to original sources

A guide to choosing data compression methods for cosmological inference

We provide a pedagogical guide to help researchers choose an appropriate compression method for cosmological inference and numerical covariance estimation, trading off the balance between information loss and complexity. We describe methods used in the literature, categorising them by the form of the loss function -- Fisher information, mutual information, or mean squared error -- and by whether they are linear or non-linear. We consider in detail the linear compression methods: massively optimised parameter estimation and data compression (MOPED), canonical correlation analysis (CCA), principal component analysis, and linear neural networks. We also investigate simple non-linear methods based on neural networks with similar architectures but different loss functions. We use tomographic weak lensing power spectra as example data vectors, deriving constraints on the cosmological parameters $\Omega_\mathrm{m}$, $\sigma_8$, $w_0$ and $w_a$ from compressed data. For our figure of merit (FoM) we use the log determinant of the Fisher matrix, which approximates the covariance of the full posterior well. MOPED and a Fisher information-based linear neural network both achieve lossless compression. Neural networks with loss functions which optimise mutual information achieve at most 84 per cent of the uncompressed FoM and those which minimise mean squared errors achieve at most 72 per cent. We show how to determine whether MOPED will retain information even if the fiducial cosmology is not known with certainty, and demonstrate a strategy for CCA which results in an FoM close to that of MOPED while requiring only an approximate idea of where the bulk of the posterior mass is located.

astro-ph.CO

Field-level weak lensing cosmology with $60$ simulations using multifidelity simulation-based inference

We perform a realistic KiDS-Legacy mock analysis with field-level neural compression and simulation-based inference using just 60 $N$-body simulations. The weak lensing shear field encodes substantially more cosmological information than standard two-point summary statistics such as the power spectrum. Field-level inference can fully exploit this information, but physical realism at the field-level requires very high-fidelity simulations. This poses a major challenge for simulation-based inference (SBI): accurate empirical density modelling and deep-learning-based neural compression require tens of thousands of training samples, but achieving physical realism at the field level makes each simulation extremely costly. We demonstrate that multifidelity SBI can alleviate this tension by substantially reducing the number of high-fidelity simulations needed for accurate cosmological inference. We pre-train neural inference models on realistic KiDS-Legacy-like shear mocks using fast log-normal \texttt{GLASS} simulations and fine-tune them on a small set of high-fidelity $N$-body simulations. We show that $60$ high-fidelity simulations are sufficient to obtain informative and well-calibrated cosmological posteriors, enabling at least an order-of-magnitude reduction in simulation cost for accurate field-level inference in a realistic setting.

astro-ph.CO

Savage-Dickey density ratio estimation with normalizing flows for Bayesian model comparison

A core motivation of science is to evaluate which scientific model best explains observed data. Bayesian model comparison provides a principled statistical approach to comparing scientific models and has found widespread application within cosmology and astrophysics. Calculating the Bayesian evidence is computationally challenging, especially as we continue to explore increasingly more complex models. The Savage-Dickey density ratio (SDDR) provides a method to calculate the Bayes factor (evidence ratio) between two nested models using only posterior samples from the super model. The SDDR requires the calculation of a normalised marginal distribution over the extra parameters of the super model, which has typically been performed using classical density estimators, such as histograms. Classical density estimators, however, can struggle to scale to high-dimensional settings. We introduce a neural SDDR approach using normalizing flows that can scale to settings where the super model contains a large number of extra parameters. We demonstrate the effectiveness of this neural SDDR methodology applied to both toy and realistic cosmological examples. For a field-level inference setting, we show that Bayes factors computed for a Bayesian hierarchical model (BHM) and simulation-based inference (SBI) approach are consistent, providing further validation that SBI extracts as much cosmological information from the field as the BHM approach. The SDDR estimator with normalizing flows is implemented in the open-source harmonic Python package.

astro-ph.CO

Simulation-based inference with scattering representations: scattering is all you need

We demonstrate the successful use of scattering representations without further compression for simulation-based inference (SBI) with images (i.e. field-level), illustrated with a cosmological case study. Scattering representations provide a highly effective representational space for subsequent learning tasks, although the higher dimensional compressed space introduces challenges. We overcome these through spatial averaging, coupled with more expressive density estimators. Compared to alternative methods, such an approach does not require additional simulations for either training or computing derivatives, is interpretable, and resilient to covariate shift. As expected, we show that a scattering only approach extracts more information than traditional second order summary statistics.

cs.LG

KiDS-SBI: Simulation-based inference analysis of KiDS-1000 cosmic shear

We present a simulation-based inference (SBI) cosmological analysis of cosmic shear two-point statistics from the fourth weak gravitational lensing data release of the ESO Kilo-Degree Survey (KiDS-1000). KiDS-SBI efficiently performs non-Limber projection of the matter power spectrum via Levin's method, and constructs log-normal random matter fields on the curved sky for arbitrary cosmologies, including effective prescriptions for intrinsic alignments and baryonic feedback. The forward model samples realistic galaxy positions and shapes based on the observational characteristics, incorporating shear measurement and redshift calibration uncertainties, as well as angular anisotropies due to variations in depth and point-spread function. To enable direct comparison with standard inference, we limit our analysis to pseudo-angular power spectra. The SBI is based on sequential neural likelihood estimation to infer the posterior distribution of spatially-flat $\Lambda$CDM cosmological parameters from 18,000 realisations. We infer a mean marginal of the growth of structure parameter $S_{8} \equiv \sigma_8 (\Omega_\mathrm{m} / 0.3)^{0.5} = 0.731\pm 0.033$ ($68 \%$). We present a measure of goodness-of-fit for SBI and determine that the forward model fits the data well with a probability-to-exceed of $0.42$. For fixed cosmology, the learnt likelihood is approximately Gaussian, while constraints widen compared to a Gaussian likelihood analysis due to cosmology dependence in the covariance. Neglecting variable depth and anisotropies in the point spread function in the model can cause $S_{8}$ to be overestimated by ${\sim}5\%$. Our results are in agreement with previous analysis of KiDS-1000 and reinforce a $2.9 \sigma$ tension with constraints from cosmic microwave background measurements. This work highlights the importance of forward-modelling systematic effects in upcoming galaxy surveys.

astro-ph.CO

A simulation-based inference pipeline for cosmic shear with the Kilo-Degree Survey

The standard approach to inference from cosmic large-scale structure data employs summary statistics that are compared to analytic models in a Gaussian likelihood with pre-computed covariance. To overcome the idealising assumptions about the form of the likelihood and the complexity of the data inherent to the standard approach, we investigate simulation-based inference (SBI), which learns the likelihood as a probability density parameterised by a neural network. We construct suites of simulated, exactly Gaussian-distributed data vectors for the most recent Kilo-Degree Survey (KiDS) weak gravitational lensing analysis and demonstrate that SBI recovers the full 12-dimensional KiDS posterior distribution with just under $10^4$ simulations. We optimise the simulation strategy by initially covering the parameter space by a hypercube, followed by batches of actively learnt additional points. The data compression in our SBI implementation is robust to suboptimal choices of fiducial parameter values and of data covariance. Together with a fast simulator, SBI is therefore a competitive and more versatile alternative to standard inference.

astro-ph.CO

Information Mechanics

Despite the wide usage of information as a concept in science, we have yet to develop a clear & concise scientific definition. This paper is aimed at laying the foundations for a new theory concerning the mechanics of information alongside its intimate relationship with working processes. Principally it aims to provide a better understanding of what information is. We find that like entropy, information is also a state variable, and both their values are surprisingly contextual. Conversely, contrary to popular belief, we find that information is not negative entropy. However, unlike entropy, information can be both positive and negative. We further find that it is possible to consider a communications process as a working process and that Shannon's entropy is only applicable to the modelling of statistical distributions. In extension, it appears that information could exist within any system, even thermodynamic systems at equilibrium. Surprisingly, the amount of mechanical work a thermodynamic system can do directly relates to the appropriate corresponding information available. If the system does not have corresponding information, it can do no work irrespective of how much energy it may contain. Following the theory we present, it will become evident to the reader that information is an intrinsic property which exists naturally within our universe, and not an abstract human notion.

physics.gen-ph