Searcharxiv⌕ Search

arXiv subjects

Chirag Modi

Publications and source records attributed to Chirag Modi.

At least 55 records · Page 3Linked to original sources

${\rm S{\scriptsize IM}BIG}$: Mock Challenge for a Forward Modeling Approach to Galaxy Clustering

Simulation-Based Inference of Galaxies (${\rm S{\scriptsize IM}BIG}$) is a forward modeling framework for analyzing galaxy clustering using simulation-based inference. In this work, we present the ${\rm S{\scriptsize IM}BIG}$ forward model, which is designed to match the observed SDSS-III BOSS CMASS galaxy sample. The forward model is based on high-resolution ${\rm Q{\scriptsize UIJOTE}}$ $N$-body simulations and a flexible halo occupation model. It includes full survey realism and models observational systematics such as angular masking and fiber collisions. We present the "mock challenge" for validating the accuracy of posteriors inferred from ${\rm S{\scriptsize IM}BIG}$ using a suite of 1,500 test simulations constructed using forward models with a different $N$-body simulation, halo finder, and halo occupation prescription. As a demonstration of ${\rm S{\scriptsize IM}BIG}$, we analyze the power spectrum multipoles out to $k_{\rm max} = 0.5\,h/{\rm Mpc}$ and infer the posterior of $Λ$CDM cosmological and halo occupation parameters. Based on the mock challenge, we find that our constraints on $Ω_m$ and $σ_8$ are unbiased, but conservative. Hence, the mock challenge demonstrates that ${\rm S{\scriptsize IM}BIG}$ provides a robust framework for inferring cosmological parameters from galaxy clustering on non-linear scales and a complete framework for handling observational systematics. In subsequent work, we will use ${\rm S{\scriptsize IM}BIG}$ to analyze summary statistics beyond the power spectrum including the bispectrum, marked power spectrum, skew spectrum, wavelet statistics, and field-level statistics.

astro-ph.CO↗

Towards a non-Gaussian Generative Model of large-scale Reionization Maps

High-dimensional data sets are expected from the next generation of large-scale surveys. These data sets will carry a wealth of information about the early stages of galaxy formation and cosmic reionization. Extracting the maximum amount of information from the these data sets remains a key challenge. Current simulations of cosmic reionization are computationally too expensive to provide enough realizations to enable testing different statistical methods, such as parameter inference. We present a non-Gaussian generative model of reionization maps that is based solely on their summary statistics. We reconstruct large-scale ionization fields (bubble spatial distributions) directly from their power spectra (PS) and Wavelet Phase Harmonics (WPH) coefficients. Using WPH, we show that our model is efficient in generating diverse new examples of large-scale ionization maps from a single realization of a summary statistic. We compare our model with the target ionization maps using the bubble size statistics, and largely find a good agreement. As compared to PS, our results show that WPH provide optimal summary statistics that capture most of information out of a highly non-linear ionization fields.

astro-ph.CO↗

Reconstructing the Universe with Variational self-Boosted Sampling

Forward modeling approaches in cosmology have made it possible to reconstruct the initial conditions at the beginning of the Universe from the observed survey data. However the high dimensionality of the parameter space still poses a challenge to explore the full posterior, with traditional algorithms such as Hamiltonian Monte Carlo (HMC) being computationally inefficient due to generating correlated samples and the performance of variational inference being highly dependent on the choice of divergence (loss) function. Here we develop a hybrid scheme, called variational self-boosted sampling (VBS) to mitigate the drawbacks of both these algorithms by learning a variational approximation for the proposal distribution of Monte Carlo sampling and combine it with HMC. The variational distribution is parameterized as a normalizing flow and learnt with samples generated on the fly, while proposals drawn from it reduce auto-correlation length in MCMC chains. Our normalizing flow uses Fourier space convolutions and element-wise operations to scale to high dimensions. We show that after a short initial warm-up and training phase, VBS generates better quality of samples than simple VI approaches and reduces the correlation length in the sampling phase by a factor of 10-50 over using only HMC to explore the posterior of initial conditions in 64$^3$ and 128$^3$ dimensional problems, with larger gains for high signal-to-noise data observations.

astro-ph.IM↗

The DESI $N$-body Simulation Project -- II. Suppressing sample variance with fast simulations

Dark Energy Spectroscopic Instrument (DESI) will construct a large and precise three-dimensional map of our Universe. The survey effective volume reaches $\sim20\Gpchcube$. It is a great challenge to prepare high-resolution simulations with a much larger volume for validating the DESI analysis pipelines. \textsc{AbacusSummit} is a suite of high-resolution dark-matter-only simulations designed for this purpose, with $200\Gpchcube$ (10 times DESI volume) for the base cosmology. However, further efforts need to be done to provide a more precise analysis of the data and to cover also other cosmologies. Recently, the CARPool method was proposed to use paired accurate and approximate simulations to achieve high statistical precision with a limited number of high-resolution simulations. Relying on this technique, we propose to use fast quasi-$N$-body solvers combined with accurate simulations to produce accurate summary statistics. This enables us to obtain 100 times smaller variance than the expected DESI statistical variance at the scales we are interested in, e.g. $k < 0.3\hMpc$ for the halo power spectrum. In addition, it can significantly suppress the sample variance of the halo bispectrum. We further generalize the method for other cosmologies with only one realization in \textsc{AbacusSummit} suite to extend the effective volume $\sim 20$ times. In summary, our proposed strategy of combining high-fidelity simulations with fast approximate gravity solvers and a series of variance suppression techniques sets the path for a robust cosmological analysis of galaxy survey data.

astro-ph.CO↗

Delayed rejection Hamiltonian Monte Carlo for sampling multiscale distributions

The efficiency of Hamiltonian Monte Carlo (HMC) can suffer when sampling a distribution with a wide range of length scales, because the small step sizes needed for stability in high-curvature regions are inefficient elsewhere. To address this we present a delayed rejection variant: if an initial HMC trajectory is rejected, we make one or more subsequent proposals each using a step size geometrically smaller than the last. We extend the standard delayed rejection framework by allowing the probability of a retry to depend on the probability of accepting the previous proposal. We test the scheme in several sampling tasks, including multiscale model distributions such as Neal's funnel, and statistical applications. Delayed rejection enables up to five-fold performance gains over optimally-tuned HMC, as measured by effective sample size per gradient evaluation. Even for simpler distributions, delayed rejection provides increased robustness to step size misspecification. Along the way, we provide an accessible but rigorous review of detailed balance for HMC.

stat.ML↗

Mind the gap: the power of combining photometric surveys with intensity mapping

The long wavelength modes lost to bright foregrounds in the interferometric 21-cm surveys can partially be recovered using a forward modeling approach that exploits the non-linear coupling between small and large scales induced by gravitational evolution. In this work, we build upon this approach by considering how adding external galaxy distribution data can help to fill in these modes. We consider supplementing the 21-cm data at two different redshifts with a spectroscopic sample (good radial resolution but low number density) loosely modeled on DESI-ELG at $z=1$ and a photometric sample (high number density but poor radial resolution) similar to LSST sample at $z=1$ and $z=4$ respectively. We find that both the galaxy samples are able to reconstruct the largest modes better than only using 21-cm data, with the spectroscopic sample performing significantly better than the photometric sample despite much lower number density. We demonstrate the synergies between surveys by showing that the primordial initial density field is reconstructed better with the combination of surveys than using either of them individually. Methodologically, we also explore the importance of smoothing the density field when using bias models to forward model these tracers for reconstruction.

astro-ph.CO↗

CosmicRIM : Reconstructing Early Universe by Combining Differentiable Simulations with Recurrent Inference Machines

Reconstructing the Gaussian initial conditions at the beginning of the Universe from the survey data in a forward modeling framework is a major challenge in cosmology. This requires solving a high dimensional inverse problem with an expensive, non-linear forward model: a cosmological N-body simulation. While intractable until recently, we propose to solve this inference problem using an automatically differentiable N-body solver, combined with a recurrent networks to learn the inference scheme and obtain the maximum-a-posteriori (MAP) estimate of the initial conditions of the Universe. We demonstrate using realistic cosmological observables that learnt inference is 40 times faster than traditional algorithms such as ADAM and LBFGS, which require specialized annealing schemes, and obtains solution of higher quality.

astro-ph.CO↗

FlowPM: Distributed TensorFlow Implementation of the FastPM Cosmological N-body Solver

We present FlowPM, a Particle-Mesh (PM) cosmological N-body code implemented in Mesh-TensorFlow for GPU-accelerated, distributed, and differentiable simulations. We implement and validate the accuracy of a novel multi-grid scheme based on multiresolution pyramids to compute large scale forces efficiently on distributed platforms. We explore the scaling of the simulation on large-scale supercomputers and compare it with corresponding python based PM code, finding on an average 10x speed-up in terms of wallclock time. We also demonstrate how this novel tool can be used for efficiently solving large scale cosmological inference problems, in particular reconstruction of cosmological fields in a forward model Bayesian framework with hybrid PM and neural network forward model. We provide skeleton code for these examples and the entire code is publicly available at https://github.com/modichirag/flowpm.

astro-ph.CO↗

Simulations and symmetries

We investigate the range of applicability of a model for the real-space power spectrum based on N-body dynamics and a (quadratic) Lagrangian bias expansion. This combination uses the highly accurate particle displacements that can be efficiently achieved by modern N-body methods with a symmetries-based bias expansion which describes the clustering of any tracer on large scales. We show that at low redshifts, and for moderately biased tracers, the substitution of N-body-determined dynamics improves over an equivalent model using perturbation theory by more than a factor of two in scale, while at high redshifts and for highly biased tracers the gains are more modest. This hybrid approach lends itself well to emulation. By removing the need to identify halos and subhalos, and by not requiring any galaxy-formation-related parameters to be included, the emulation task is significantly simplified at the cost of modeling a more limited range in scale.

astro-ph.CO↗

Intensity mapping with neutral hydrogen and the Hidden Valley simulations

This paper introduces the Hidden Valley simulations, a set of trillion-particle N-body simulations in gigaparsec volumes aimed at intensity mapping science. We present details of the simulations and their convergence, then specialize to the study of 21-cm fluctuations between redshifts 2 and 6. Neutral hydrogen is assigned to halos using three prescriptions, and we investigate the clustering in real and redshift-space at the 2-point level. In common with earlier work we find the bias of HI increases from near 2 at z = 2 to 4 at z = 6, becoming more scale dependent at high z. The level of scale-dependence and decorrelation with the matter field are as predicted by perturbation theory. Due to the low mass of the hosting halos, the impact of fingers of god is small on the range relevant for proposed 21-cm instruments. We show that baryon acoustic oscillations and redshift-space distortions could be well measured by such instruments. Taking advantage of the large simulation volume, we assess the impact of fluctuations in the ultraviolet background, which change HI clustering primarily at large scales.

astro-ph.CO↗

Reconstructing large-scale structure with neutral hydrogen surveys

Upcoming 21-cm intensity surveys will use the hyperfine transition in emission to map out neutral hydrogen in large volumes of the universe. Unfortunately, large spatial scales are completely contaminated with spectrally smooth astrophysical foregrounds which are orders of magnitude brighter than the signal. This contamination also leaks into smaller radial and angular modes to form a foreground wedge, further limiting the usefulness of 21-cm observations for different science cases, especially cross-correlations with tracers that have wide kernels in the radial direction. In this paper, we investigate reconstructing these modes within a forward modeling framework. Starting with an initial density field, a suitable bias parameterization and non-linear dynamics to model the observed 21-cm field, our reconstruction proceeds by combining the likelihood of a forward simulation to match the observations (under given modeling error and a data noise model) with the Gaussian prior on initial conditions and maximizing the obtained posterior. For redshifts $z=2$ and $4$, we are able to reconstruct 21cm field with cross correlation, $r_c > 0.8$ on all scales for both our optimistic and pessimistic assumptions about foreground contamination and for different levels of thermal noise. The performance deteriorates slightly at $z=6$. The large-scale line-of-sight modes are reconstructed almost perfectly. We demonstrate how our method also reconstructs baryon acoustic oscillations, outperforming standard methods on all scales. We also describe how our reconstructed field can provide superb clustering redshift estimation at high redshifts, where it is otherwise extremely difficult to obtain dense spectroscopic samples, as well as open up cross-correlation opportunities with projected fields (e.g. lensing) which are restricted to modes transverse to the line of sight.

astro-ph.CO↗

Lensing corrections on galaxy-lensing cross correlations and galaxy-galaxy auto correlations

We study the impact of lensing corrections on modeling cross correlations between CMB lensing and galaxies, cosmic shear and galaxies, and galaxies in different redshift bins. Estimating the importance of these corrections becomes necessary in the light of anticipated high-accuracy measurements of these observables. While higher order lensing corrections (sometimes also referred to as post Born corrections) have been shown to be negligibly small for lensing auto correlations, they have not been studied for cross correlations. We evaluate the contributing four-point functions without making use of the Limber approximation and compute line-of-sight integrals with the numerically stable and fast FFTlog formalism. We find that the relative size of lensing corrections depends on the respective redshift distributions of the lensing sources and galaxies, but that they are generally small for high signal-to-noise correlations. We point out that a full assessment and judgement of the importance of these corrections requires the inclusion of lensing Jacobian terms on the galaxy side. We identify these additional correction terms, but do not evaluate them due to their large number. We argue that they could be potentially important and suggest that their size should be measured in the future with ray-traced simulations. We make our code publicly available.

astro-ph.CO↗

Generative Learning of Counterfactual for Synthetic Control Applications in Econometrics

A common statistical problem in econometrics is to estimate the impact of a treatment on a treated unit given a control sample with untreated outcomes. Here we develop a generative learning approach to this problem, learning the probability distribution of the data, which can be used for downstream tasks such as post-treatment counterfactual prediction and hypothesis testing. We use control samples to transform the data to a Gaussian and homoschedastic form and then perform Gaussian process analysis in Fourier space, evaluating the optimal Gaussian kernel via non-parametric power spectrum estimation. We combine this Gaussian prior with the data likelihood given by the pre-treatment data of the single unit, to obtain the synthetic prediction of the unit post-treatment, which minimizes the error variance of synthetic prediction. Given the generative model the minimum variance counterfactual is unique, and comes with an associated error covariance matrix. We extend this basic formalism to include correlations of primary variable with other covariates of interest. Given the probabilistic description of generative model we can compare synthetic data prediction with real data to address the question of whether the treatment had a statistically significant impact. For this purpose we develop a hypothesis testing approach and evaluate the Bayes factor. We apply the method to the well studied example of California (CA) tobacco sales tax of 1988. We also perform a placebo analysis using control states to validate our methodology. Our hypothesis testing method suggests 5.8:1 odds in favor of CA tobacco sales tax having an impact on the tobacco sales, a value that is at least three times higher than any of the 38 control states.

stat.ML↗

Cosmological Reconstruction From Galaxy Light: Neural Network Based Light-Matter Connection

We present a method to reconstruct the initial conditions of the universe using observed galaxy positions and luminosities under the assumption that the luminosities can be calibrated with weak lensing to give the mean halo mass. Our method relies on following the gradients of forward model and since the standard way to identify halos is non-differentiable and results in a discrete sample of objects, we propose a framework to model the halo position and mass field starting from the non-linear matter field using Neural Networks. We evaluate the performance of our model with multiple metrics. Our model is more than $95\%$ correlated with the halo-mass fields up to $k\sim 0.7 {\rm h/Mpc}$ and significantly reduces the stochasticity over the Poisson shot noise. We develop a data likelihood model that takes our modeling error and intrinsic scatter in the halo mass-light relation into account and show that a displaced log-normal model is a good approximation to it. We optimize over the corresponding loss function to reconstruct the initial density field and develop an annealing procedure to speed up and improve the convergence. We apply the method to halo number densities of $\bar{n} = 2.5\times 10^{-4} -10^{-3}({\rm h/Mpc})^3$, typical of current and future redshift surveys, and recover a Gaussian initial density field, mapping all the higher order information in the data into the power spectrum. We show that our reconstruction improves over the standard reconstruction. For baryonic acoustic oscillations (BAO) the gains are relatively modest because BAO is dominated by large scales where standard reconstruction suffices. We improve upon it by $\sim 15-20\%$ in terms of error on BAO peak as estimated by Fisher analysis at $z=0$. We expect larger gains will be achieved when applying this method to the broadband linear power spectrum reconstruction on smaller scales.

astro-ph.CO↗

Towards optimal extraction of cosmological information from nonlinear data

One of the main unsolved problems of cosmology is how to maximize the extraction of information from nonlinear data. If the data are nonlinear the usual approach is to employ a sequence of statistics (N-point statistics, counting statistics of clusters, density peaks or voids etc.), along with the corresponding covariance matrices. However, this approach is computationally prohibitive and has not been shown to be exhaustive in terms of information content. Here we instead develop a Bayesian approach, expanding the likelihood around the maximum posterior of linear modes, which we solve for using optimization methods. By integrating out the modes using perturbative expansion of the likelihood we construct an initial power spectrum estimator, which for a fixed forward model contains all the cosmological information if the initial modes are gaussian distributed. We develop a method to construct the window and covariance matrix such that the estimator is explicitly unbiased and nearly optimal. We then generalize the method to include the forward model parameters, including cosmological and nuisance parameters, and primordial non-gaussianity. We apply the method in the simplified context of nonlinear structure formation, using either simplified 2-LPT dynamics or N-body simulations as the nonlinear mapping between linear and nonlinear density, and 2-LPT dynamics in the optimization steps used to reconstruct the initial density modes. We demonstrate that the method gives an unbiased estimator of the initial power spectrum, providing among other a near optimal reconstruction of linear baryonic acoustic oscillations.

astro-ph.CO↗

nbodykit: an open-source, massively parallel toolkit for large-scale structure

We present nbodykit, an open-source, massively parallel Python toolkit for analyzing large-scale structure (LSS) data. Using Python bindings of the Message Passing Interface (MPI), we provide parallel implementations of many commonly used algorithms in LSS. nbodykit is both an interactive and scalable piece of scientific software, performing well in a supercomputing environment while still taking advantage of the interactive tools provided by the Python ecosystem. Existing functionality includes estimators of the power spectrum, 2 and 3-point correlation functions, a Friends-of-Friends grouping algorithm, mock catalog creation via the halo occupation distribution technique, and approximate N-body simulations via the FastPM scheme. The package also provides a set of distributed data containers, insulated from the algorithms themselves, that enable nbodykit to provide a unified treatment of both simulation and observational data sets. nbodykit can be easily deployed in a high performance computing environment, overcoming some of the traditional difficulties of using Python on supercomputers. We provide performance benchmarks illustrating the scalability of the software. The modular, component-based approach of nbodykit allows researchers to easily build complex applications using its tools. The package is extensively documented at http://nbodykit.readthedocs.io, which also includes an interactive set of example recipes for new users to explore. As open-source software, we hope nbodykit provides a common framework for the community to use and develop in confronting the analysis challenges of future LSS surveys.

astro-ph.IM↗

Halo bias in Lagrangian Space: Estimators and theoretical predictions

We present several methods to accurately estimate Lagrangian bias parameters and substantiate them using simulations. In particular, we focus on the quadratic terms, both the local and the non local ones, and show the first clear evidence for the latter in the simulations. Using Fourier space correlations, we also show for the first time, the scale dependence of the quadratic and non-local bias coefficients. For the linear bias, we fit for the scale dependence and demonstrate the validity of a consistency relation between linear bias parameters. Furthermore we employ real space estimators, using both cross-correlations and the Peak-Background Split argument. This is the first time the latter is used to measure anisotropic bias coefficients. We find good agreement for all the parameters among these different methods, and also good agreement for local bias with ESP$τ$ theory predictions. We also try to exploit possible relations among the different bias parameters. Finally, we show how including higher order bias reduces the magnitude and scale dependence of stochasticity of the halo field.

astro-ph.CO↗

Modeling CMB Lensing Cross Correlations with {\sc CLEFT}

A new generation of surveys will soon map large fractions of sky to ever greater depths and their science goals can be enhanced by exploiting cross correlations between them. In this paper we study cross correlations between the lensing of the CMB and biased tracers of large-scale structure at high $z$. We motivate the need for more sophisticated bias models for modeling increasingly biased tracers at these redshifts and propose the use of perturbation theories, specifically Convolution Lagrangian Effective Field Theory ({\sc CLEFT}). Since such signals reside at large scales and redshifts, they can be well described by perturbative approaches. We compare our model with the current approach of using scale independent bias coupled with fitting functions for non-linear matter power spectra, showing that the latter will not be sufficient for upcoming surveys. We illustrate our ideas by estimating $σ_8$ from the auto- and cross-spectra of mock surveys, finding that {\sc CLEFT} returns accurate and unbiased results at high $z$. We discuss uncertainties due to the redshift distribution of the tracers, and several avenues for future development.

astro-ph.CO↗