SearcharxivSearch

arXiv subjects

Tomasz Kacprzak

Publications and source records attributed to Tomasz Kacprzak.

At least 19 recordsLinked to original sources

GalSBI: Forward Modelling Galaxy Clustering and Population

Forward modelling is a powerful approach for analyzing large-scale structure surveys. For this purpose, we extend the GalSBI framework to jointly model the galaxy population and clustering using an efficient subhalo abundance matching scheme based on optimal transport. We use simulation-based inference to constrain the model parameters by comparing UFig image simulations with DES Y3 imaging data. As a validation, we find that galaxy photometry and morphology agree well with multi-band imaging data of different depths, namely DES and HSC deep fields. Galaxy clustering for simulation and data is also in good agreement when comparing the angular power spectrum for different magnitude and color cuts. We further compare simulated redshift distributions against high-precision photometric redshifts in HSC deep field imaging of the COSMOS field. We find the redshift distributions across magnitude cuts to be similar to previous work, however with more realistic uncertainty modelling due to the addition of clustering contribution to sample variance. The agreement of the mean redshifts with data is very good, between $0.2σ$ and $1.6σ$ for different magnitude cuts, with sample variance being the dominant uncertainty contributor in bright samples ($<24$ mag) and subdominant compared to galaxy population model uncertainty in fainter samples. As a byproduct we measure the galaxy luminosity function and galaxy-halo connection, which are broadly consistent with existing literature. The updated GalSBI code and galaxy population model are publicly available. They enable accurate forward-modelled image simulations with realistic clustering, which can be used to model the effect of sample variance, source clustering, redshift distributions, and blending in large-scale-structure surveys. This makes GalSBI a powerful tool for the analysis of current and next-generation cosmological galaxy surveys.

astro-ph.CO

Increasing the Precision of Surrogate Models for Weak Lensing Mass Maps with Flow Matching

Weak gravitational lensing maps compactly encode the evolution of cosmic large-scale structure and are a key tool for cosmological analyses. Performing inference directly at the map level allows flexible choices of statistics and can increase constraining power. Conventional methods rely solely on N-body simulations and are computationally expensive. Generative machine-learning emulators can accelerate map-level theory prediction. However, existing GAN-based map-level surrogates still have limited statistical fidelity. They can produce over-smoothed maps, may fail to capture the full distribution of generated map sets and can be difficult to train. Continuous normalizing flows trained with flow matching have recently emerged as a powerful class of generative models. We present a residual label-conditional flow matching generative network that conditions explicitly on the matter density Omega_m and clustering amplitude sigma_8 for a fixed source redshift distribution n(z). The model learns a continuous probability flow in a residual space from label-specific noise distributions to convergence maps. We evaluate it using pixel and peak statistics, the power spectrum, bispectrum, power-spectrum correlation matrices, and other validation metrics. Compared with the previous GAN benchmark, the proposed method improves the typical fidelity of generated maps from below 10% and below 20% to below 1% and below 5% for basic and higher-order statistics, respectively. The agreement at the level of map distributions is also very good: maps generated from random noise match well the distribution of maps generated with N-body simulations from random initial conditions. This work brings us closer to a practical mass-map emulator that captures the cosmological signal while supporting multiple forms of data analysis.

astro-ph.CO

galsbi: A Python package for the GalSBI galaxy population model

Large-scale structure surveys measure the shapes and positions of millions of galaxies in order to constrain the cosmological model with high precision. The resulting large data volume poses a challenge for the analysis of the data, from the estimation of photometric redshifts to the calibration of shape measurements. We present GalSBI, a model for the galaxy population, to address these challenges. This phenomenological model is constrained by observational data using simulation-based inference (SBI). The $\texttt{galsbi}$ Python package provides an easy interface to generate catalogs of galaxies based on the GalSBI model, including their photometric properties, and to simulate realistic images of these galaxies using the $\texttt{UFig}$ package.

astro-ph.CO

GalSBI-SPS: a stellar population synthesis-based galaxy population model for cosmology and galaxy evolution applications

Next generation photometric and spectroscopic surveys will enable unprecedented tests of the concordance cosmological model and of galaxy formation and evolution. Fully exploiting their potential requires a precise understanding of the selection effects on galaxies and biases on measurements of their properties, required, above all, for accurate estimates of redshift distributions n(z). Forward-modelling offers a powerful framework to simultaneously recover galaxy $n(z)$s and characterise the observed galaxy population. We present GalSBI-SPS, a new SPS-based galaxy population model that generates realistic galaxy catalogues, which we use to forward-model HSC data in the COSMOS field. GalSBI-SPS samples galaxy physical properties, computes magnitudes with ProSpect, and simulates HSC images in the COSMOS field with UFig. We measure photometric properties consistently in real data and simulations. We compare $n(z)$s, photometric and physical properties to observations and to GalSBI. GalSBI-SPS reproduces the observed grizy magnitude, colour, and size distributions down to i<23. Median differences in magnitudes and colours remain below 0.14 mag, with the model covering the full colour space spanned by HSC. Galaxy sizes are overestimated by 0.2 arcsec on average and some tension exists in the g-r colour, but the latter is comparable to that seen in GalSBI. $n(z)$s show a mild positive offset (0.01-0.08) in the mean. GalSBI-SPS qualitatively reproduces the stellar mass-SFR and size-stellar mass relations seen in COSMOS2020. GalSBI-SPS provides a realistic, survey-independent galaxy population description at a Stage-III depth using only literature-based parameters. Its predictive power will improve significantly when constrained against observed data using SBI, thereby providing accurate $n(z)$s satisfying the stringent requirements set by Stage IV surveys.

astro-ph.GA

UFig v1: The ultra-fast image generator

With the rise of simulation-based inference (SBI) methods, simulations need to be fast as well as realistic. $\texttt{UFig v1}$ is a public Python package that simulates astronomical images with exceptional speed, taking approximately the same time as source extraction. This makes it particularly well-suited for SBI methods where computational efficiency is crucial. To render an image, $\texttt{UFig}$ requires a galaxy catalog, and a description of the point spread function (PSF). It can also add background noise, sample stars using the Besançon model of the Milky Way, and run $\texttt{SExtractor}$ to extract sources from the rendered image. The extracted sources can be matched to the intrinsic catalog, flagged based on $\texttt{SExtractor}$ output and survey masks, and emulators can be used to bypass the image simulation and extraction steps. A first version of $\texttt{UFig}$ was presented in Bergé et al. (2013) and the software has since been used and further developed in a variety of forward modelling applications.

astro-ph.IM

SHAM-OT: Rapid Subhalo Abundance Matching with Optimal Transport

Subhalo abundance matching (SHAM) is widely used for connecting galaxies to dark matter haloes. In SHAM, galaxies and (sub-)haloes are sorted according to their mass (or mass proxy) and matched by their rank order. In this work, we show that SHAM is the solution of the optimal transport (OT) problem on empirical distributions (samples or catalogues) for any metric transport cost function. In the limit of large number of samples, it converges to the solution of the OT problem between continuous distributions. We propose SHAM-OT: a formulation of abundance matching where the halo-galaxy relation is obtained as the optimal transport plan between galaxy and halo mass functions. By working directly on these (discretized) functions, SHAM-OT eliminates the need for sampling or sorting and is solved using efficient OT algorithms at negligible compute and memory cost. Scatter in the galaxy-halo relation can be naturally incorporated through regularization of the transport plan. SHAM-OT can easily be generalized to multiple marginal distributions. We validate our method using analytical tests with varying cosmology and luminosity function parameters, and on simulated halo catalogues. The efficiency of SHAM-OT makes it particularly advantageous for Bayesian inference that requires marginalization over stellar mass or luminosity function uncertainties.

astro-ph.CO

GalSBI: Phenomenological galaxy population model for cosmology using simulation-based inference

We present GalSBI, a phenomenological model of the galaxy population for cosmological applications using simulation-based inference. The model is based on analytical parametrizations of galaxy luminosity functions, morphologies and spectral energy distributions. Model constraints are derived through iterative Approximate Bayesian Computation, by comparing Hyper Suprime-Cam deep field images with simulations which include a forward model of instrumental, observational and source extraction effects. We developed an emulator trained on image simulations using a normalizing flow. We use it to accelerate the inference by predicting detection probabilities, including blending effects and photometric properties of each object, while accounting for background and PSF variations. This enables robustness tests for all elements of the forward model and the inference. The model demonstrates excellent performance when comparing photometric properties from simulations with observed imaging data for key parameters such as magnitudes, colors and sizes. The redshift distribution of simulated galaxies agrees well with high-precision photometric redshifts in the COSMOS field within $1.5σ$ for all magnitude cuts. Additionally, we demonstrate how GalSBI's redshifts can be utilized for splitting galaxy catalogs into tomographic bins, highlighting its potential for current and upcoming surveys. GalSBI is fully open-source, with the accompanying Python package, $\texttt{galsbi}$, offering an easy interface to quickly generate realistic, survey-independent galaxy catalogs.

astro-ph.CO

Scalable Approximate Algorithms for Optimal Transport Linear Models

Recently, linear regression models incorporating an optimal transport (OT) loss have been explored for applications such as supervised unmixing of spectra, music transcription, and mass spectrometry. However, these task-specific approaches often do not generalize readily to a broader class of linear models. In this work, we propose a novel algorithmic framework for solving a general class of non-negative linear regression models with an entropy-regularized OT datafit term, based on Sinkhorn-like scaling iterations. Our framework accommodates convex penalty functions on the weights (e.g. squared-$\ell_2$ and $\ell_1$ norms), and admits additional convex loss terms between the transported marginal and target distribution (e.g. squared error or total variation). We derive simple multiplicative updates for common penalty and datafit terms. This method is suitable for large-scale problems due to its simplicity of implementation and straightforward parallelization.

stat.ML

Interpretability of deep-learning methods applied to large-scale structure surveys

Deep learning and convolutional neural networks in particular are powerful and promising tools for cosmological analysis of large-scale structure surveys. They are already providing similar performance to classical analysis methods using fixed summary statistics, are showing potential to break key degeneracies by better probe combination and will likely improve rapidly in the coming years as progress is made in the physical modelling through both software and hardware improvement. One key issue remains: unlike classical analysis, a convolutional neural network's decision process is hidden from the user as the network optimises millions of parameters with no direct physical meaning. This prevents a clear understanding of the potential limitations and biases of the analysis, making it hard to rely on as a main analysis method. In this work, we explore the behaviour of such a convolutional neural network through a novel method. Instead of trying to analyse a network a posteriori, i.e. after training has been completed, we study the impact on the constraining power of training the network and predicting parameters with degraded data where we removed part of the information. This allows us to gain an understanding of which parts and features of a large-scale structure survey are most important in the network's prediction process. We find that the network's prediction process relies on a mix of both Gaussian and non-Gaussian information, and seems to put an emphasis on structures whose scales are at the limit between linear and non-linear regimes.

astro-ph.CO

Laue Indexing with Optimal Transport

Laue tomography experiments retrieve the positions and orientations of crystal grains in a polycrystalline samples from diffraction patterns recorded at multiple viewing angles. The use of a broad wavelength spectrum beam can greatly reduce the experimental time, but poses a difficult challenge for the indexing of diffraction peaks in polycrystalline samples; the information about the wavelength of these Bragg peaks is absent and the diffraction patterns from multiple grains are superimposed. To date, no algorithms exist capable of indexing samples with more than about 500 grains efficiently. To address this need we present a novel method: Laue indexing with Optimal Transport (LaueOT). We create a probabilistic description of the multi-grain indexing problem and propose a solution based on Sinkhorn Expectation-Maximization method, which allows to efficiently find the maximum of the likelihood thanks to the assignments being calculated using Optimal Transport. This is a non-convex optimization problem, where the orientations and positions of grains are optimized simultaneously with grain-to-spot assignments, while robustly handling the outliers. The selection of initial prototype grains to consider in the optimization problem are also calculated within the Optimal Transport framework. LaueOT can rapidly and effectively index up to 1000 grains on a single large memory GPU within less than 30 minutes. We demonstrate the performance of LaueOT on simulations with variable numbers of grains, spot position measurement noise levels, and outlier fractions. The algorithm recovers the correct number of grains even for high noise levels and up to 70% outliers in our experiments. We compare the results of indexing with LaueOT to existing algorithms both on synthetic and real neutron diffraction data from well-characterized samples.

cond-mat.mtrl-sci

Simulation-based inference of deep fields: galaxy population model and redshift distributions

Accurate redshift calibration is required to obtain unbiased cosmological information from large-scale galaxy surveys. In a forward modelling approach, the redshift distribution n(z) of a galaxy sample is measured using a parametric galaxy population model constrained by observations. We use a model that captures the redshift evolution of the galaxy luminosity functions, colours, and morphology, for red and blue samples. We constrain this model via simulation-based inference, using factorized Approximate Bayesian Computation (ABC) at the image level. We apply this framework to HSC deep field images, complemented with photometric redshifts from COSMOS2020. The simulated telescope images include realistic observational and instrumental effects. By applying the same processing and selection to real data and simulations, we obtain a sample of n(z) distributions from the ABC posterior. The photometric properties of the simulated galaxies are in good agreement with those from the real data, including magnitude, colour and redshift joint distributions. We compare the posterior n(z) from our simulations to the COSMOS2020 redshift distributions obtained via template fitting photometric data spanning the wavelength range from UV to IR. We mitigate sample variance in COSMOS by applying a reweighting technique. We thus obtain a good agreement between the simulated and observed redshift distributions, with a difference in the mean at the 1$σ$ level up to a magnitude of 24 in the i band. We discuss how our forward model can be applied to current and future surveys and be further extended. The ABC posterior and further material will be made publicly available at https://cosmology.ethz.ch/research/software-lab/ufig.html.

astro-ph.CO

Fast Forward Modelling of Galaxy Spatial and Statistical Distributions

A forward modelling approach provides simple, fast and realistic simulations of galaxy surveys, without a complex underlying model. For this purpose, galaxy clustering needs to be simulated accurately, both for the usage of clustering as its own probe and to control systematics. We present a forward model to simulate galaxy surveys, where we extend the Ultra-Fast Image Generator to include galaxy clustering. We use the distribution functions of the galaxy properties, derived from a forward model adjusted to observations. This population model jointly describes the luminosity functions, sizes, ellipticities, SEDs and apparent magnitudes. To simulate the positions of galaxies, we then use a two-parameter relation between galaxies and halos with Subhalo Abundance Matching (SHAM). We simulate the halos and subhalos using the fast PINOCCHIO code, and a method to extract the surviving subhalos from the merger history. Our simulations contain a red and a blue galaxy population, for which we build a SHAM model based on star formation quenching. For central galaxies, mass quenching is controlled with the parameter M$_{\mathrm{limit}}$, with blue galaxies residing in smaller halos. For satellite galaxies, environmental quenching is implemented with the parameter t$_{\mathrm{quench}}$, where blue galaxies occupy only recently merged subhalos. We build and test our model by comparing to imaging data from the Dark Energy Survey Year 1. To ensure completeness in our simulations, we consider the brightest galaxies with $i<20$. We find statistical agreement between our simulations and the data for two-point correlation functions on medium to large scales. Our model provides constraints on the two SHAM parameters M$_{\mathrm{limit}}$ and t$_{\mathrm{quench}}$ and offers great prospects for the quick generation of galaxy mock catalogues, optimized to agree with observations.

astro-ph.GA

$\mathbf{12\times2}$pt combined probes: pipeline, neutrino mass, and data compression

With the rapid advance of wide-field surveys it is increasingly important to perform combined cosmological probe analyses. We present a new pipeline for simulation-based multi-probe analyses, which combines tomographic large-scale structure (LSS) probes (weak lensing and galaxy clustering) with cosmic microwave background (CMB) primary and lensing data. These are combined at the $C_\ell$-level, yielding 12 distinct auto- and cross-correlations. The pipeline is based on $\texttt{UFalconv2}$, a framework to generate fast, self-consistent map-level realizations of cosmological probes from input lightcones, which is applied to the $\texttt{CosmoGridV1}$ N-body simulation suite. It includes a non-Gaussian simulation-based covariance for the LSS tracers, several data compression schemes, and a neural network emulator for accelerated theoretical predictions. We validate our framework, apply it to a simulated $12\times2$pt tomographic analysis of KiDS, BOSS, and $\textit{Planck}$, and forecast constraints for a $Λ$CDM model with a variable neutrino mass. We find that, while the neutrino mass constraints are driven by the CMB data, the addition of LSS data helps to break degeneracies and improves the constraint by up to 35%. For a fiducial $M_ν=0.15\mathrm{eV}$, a full combination of the above CMB+LSS data would enable a $3σ$ constraint on the neutrino mass. We explore data compression schemes and find that MOPED outperforms PCA. We also study the impact of an internal lensing tension in the CMB data, parametrized by $A_L$, on the neutrino mass constraint, finding that the addition of LSS to CMB data including all cross-correlations is able to mitigate the impact of this systematic. $\texttt{UFalconv2}$ and a MOPED compressed $\textit{Planck}$ CMB primary + CMB lensing likelihood are made publicly available. [abridged]

astro-ph.CO

Towards a full $w$CDM map-based analysis for weak lensing surveys

The next generation of weak lensing surveys will measure the matter distribution of the local Universe with unprecedented precision, allowing the resolution of non-Gaussian features of the convergence field. This encourages the use of higher-order mass-map statistics for cosmological parameter inference. We extend the forward-modelling based methodology introduced in a previous forecast paper to match these new requirements. We provide multiple forecasts for the wCDM parameter constraints that can be expected from stage 3 and 4 weak lensing surveys. We consider different survey setups, summary statistics and mass map filters including wavelets. We take into account the shear bias, photometric redshift uncertainties and intrinsic alignment. The impact of baryons is investigated and the necessary scale cuts are applied. We compare the angular power spectrum analysis to peak and minima counts as well as Minkowski functionals of the mass maps. We find a preference for Starlet over Gaussian filters. Our results suggest that using a survey setup with 10 instead of 5 tomographic redshift bins is beneficial. Adding cross-tomographic information improves the constraints on cosmology and especially on galaxy intrinsic alignment for all statistics. In terms of constraining power, we find the angular power spectrum and the peak counts to be equally matched for stage 4 surveys, followed by minima counts and the Minkowski functionals. Combining different summary statistics significantly improves the constraints and compensates the stringent scale cuts. We identify the most `cost-effective' combination to be the angular power spectrum, peak counts and Minkowski functionals following Starlet filtering.

astro-ph.CO

The Third Gravitational Lensing Accuracy Testing (GREAT3) Challenge Handbook

The GRavitational lEnsing Accuracy Testing 3 (GREAT3) challenge is the third in a series of image analysis challenges, with a goal of testing and facilitating the development of methods for analyzing astronomical images that will be used to measure weak gravitational lensing. This measurement requires extremely precise estimation of very small galaxy shape distortions, in the presence of far larger intrinsic galaxy shapes and distortions due to the blurring kernel caused by the atmosphere, telescope optics, and instrumental effects. The GREAT3 challenge is posed to the astronomy, machine learning, and statistics communities, and includes tests of three specific effects that are of immediate relevance to upcoming weak lensing surveys, two of which have never been tested in a community challenge before. These effects include realistically complex galaxy models based on high-resolution imaging from space; spatially varying, physically-motivated blurring kernel; and combination of multiple different exposures. To facilitate entry by people new to the field, and for use as a diagnostic tool, the simulation software for the challenge is publicly available, though the exact parameters used for the challenge are blinded. Sample scripts to analyze the challenge data using existing methods will also be provided. See http://great3challenge.info and http://great3.projects.phys.ucl.ac.uk/leaderboard/ for more information.

astro-ph.CO

GREAT3 results I: systematic errors in shear estimation and the impact of real galaxy morphology

We present first results from the third GRavitational lEnsing Accuracy Testing (GREAT3) challenge, the third in a sequence of challenges for testing methods of inferring weak gravitational lensing shear distortions from simulated galaxy images. GREAT3 was divided into experiments to test three specific questions, and included simulated space- and ground-based data with constant or cosmologically-varying shear fields. The simplest (control) experiment included parametric galaxies with a realistic distribution of signal-to-noise, size, and ellipticity, and a complex point spread function (PSF). The other experiments tested the additional impact of realistic galaxy morphology, multiple exposure imaging, and the uncertainty about a spatially-varying PSF; the last two questions will be explored in Paper II. The 24 participating teams competed to estimate lensing shears to within systematic error tolerances for upcoming Stage-IV dark energy surveys, making 1525 submissions overall. GREAT3 saw considerable variety and innovation in the types of methods applied. Several teams now meet or exceed the targets in many of the tests conducted (to within the statistical errors). We conclude that the presence of realistic galaxy morphology in simulations changes shear calibration biases by $\sim 1$ per cent for a wide range of methods. Other effects such as truncation biases due to finite galaxy postage stamps, and the impact of galaxy type as measured by the Sérsic index, are quantified for the first time. Our results generalize previous studies regarding sensitivities to galaxy size and signal-to-noise, and to PSF properties such as seeing and defocus. Almost all methods' results support the simple model in which additive shear biases depend linearly on PSF ellipticity.

astro-ph.CO

A tomographic spherical mass map emulator of the KiDS-1000 survey using conditional generative adversarial networks

Large sets of matter density simulations are becoming increasingly important in large-scale structure cosmology. Matter power spectra emulators, such as the Euclid Emulator and CosmicEmu, are trained on simulations to correct the non-linear part of the power spectrum. Map-based analyses retrieve additional non-Gaussian information from the density field, whether through human-designed statistics such as peak counts, or via machine learning methods such as convolutional neural networks. The simulations required for these methods are very resource-intensive, both in terms of computing time and storage. Map-level density field emulators, based on deep generative models, have recently been proposed to address these challenges. In this work, we present a novel mass map emulator of the KiDS-1000 survey footprint, which generates noise-free spherical maps in a fraction of a second. It takes a set of cosmological parameters $(Ω_M, σ_8)$ as input and produces a consistent set of 5 maps, corresponding to the KiDS-1000 tomographic redshift bins. To construct the emulator, we use a conditional generative adversarial network architecture and the spherical CNN $\texttt{DeepSphere}$, and train it on N-body-simulated mass maps. We compare its performance using an array of quantitative comparison metrics: angular power spectra $C_\ell$, pixel/peaks distributions, $C_\ell$ correlation matrices, and Structural Similarity Index. Overall, the average agreement on these summary statistics is $<10\%$ for the cosmologies at the centre of the simulation grid, and degrades slightly on grid edges. Finally, we perform a mock cosmological parameter estimation using the emulator and the original simulation set. We find good agreement in these constraints, for both likelihood and likelihood-free approaches. The emulator is available at https://tfhub.dev/cosmo-group-ethz/models/kids-cgan/1.

astro-ph.CO

Cosmology from Galaxy Redshift Surveys with PointNet

In recent years, deep learning approaches have achieved state-of-the-art results in the analysis of point cloud data. In cosmology, galaxy redshift surveys resemble such a permutation invariant collection of positions in space. These surveys have so far mostly been analysed with two-point statistics, such as power spectra and correlation functions. The usage of these summary statistics is best justified on large scales, where the density field is linear and Gaussian. However, in light of the increased precision expected from upcoming surveys, the analysis of -- intrinsically non-Gaussian -- small angular separations represents an appealing avenue to better constrain cosmological parameters. In this work, we aim to improve upon two-point statistics by employing a \textit{PointNet}-like neural network to regress the values of the cosmological parameters directly from point cloud data. Our implementation of PointNets can analyse inputs of $\mathcal{O}(10^4) - \mathcal{O}(10^5)$ galaxies at a time, which improves upon earlier work for this application by roughly two orders of magnitude. Additionally, we demonstrate the ability to analyse galaxy redshift survey data on the lightcone, as opposed to previously static simulation boxes at a given fixed redshift.

astro-ph.CO