SearcharxivSearch

arXiv subjects

Marco Gatti

Publications and source records attributed to Marco Gatti.

At least 19 recordsLinked to original sources

Comparing explicit likelihood and likelihood-free simulation-based inference for weak lensing cosmic shear

Simulation-based inference (SBI) has become a major tool for extracting cosmological information from weak-lensing (WL) surveys, particularly from non-Gaussian observables. We compare its two main paradigms: explicit likelihood inference (ELI), based on a Gaussian likelihood built from an emulator and covariance matrix, and likelihood-free inference (LFI), which learns the likelihood directly from simulations using neural density estimators. Using Gaussian random field mocks representative of the non-tomographic final Euclid data release, we analyse shear two-point correlation functions (shear-2PCFs), compressed with linear or non-linear methods, together with a fundamentally different map-level convolutional neural network (CNN) statistic, focusing on $\Omega_{\rm m}$ and $S_8$. We deploy posterior calibration diagnostics developed for LFI, including the test of accuracy with random points (TARP), showing that ELI becomes strongly miscalibrated under emulation inaccuracies or likelihood non-Gaussianity, whereas LFI remains well calibrated. These effects drive substantial disagreement between ELI and LFI, which largely vanishes once addressed. We further show that the compression scheme can significantly degrade ELI while leaving LFI largely unaffected. Although shear-2PCFs should capture all the information in Gaussian fields, finite compression and non-Gaussian likelihoods cause ELI constraints to differ by up to a factor of two from those inferred with the CNN, while the discrepancy drops to $\approx 30\%$ for LFI, underscoring the robustness of the deep-learning probe. Overall, our results indicate that in our simple setup, which neglects systematic biases, LFI provides a more robust and better-calibrated framework, while highlighting accurate non-Gaussian likelihood modelling and posterior calibration diagnostics as essential for future ELI analyses.

astro-ph.CO

Brightest Cluster Galaxy ellipticity as proxy for halo shape: Orientation bias, assembly bias, and potential selection effects in SZ-selected clusters

The orientation of triaxial galaxy clusters with respect to the line-of-sight is expected to be one of the prime sources of scatter and potential bias in optical observables (e.g., richness and weak-lensing signal) of galaxy clusters. In this work, we use the observed shape of the central Brightest Cluster Galaxy (BCG) as proxy for the orientation along the line-of-sight for clusters selected via the Sunyaev-Zel'dovich (SZ) effect from the South Pole Telescope (SPT) and Atacama Cosmology Telescope (ACT) surveys, matched to optically selected clusters from the Dark Energy Survey Year 3 (DES). We construct two samples of clusters that are designed to be identical in SZ mass estimate and redshift but with the roundest vs. the most elliptical BCGs, which we expect to correspond to BCGs (and clusters) with major axes aligned along the line-of-sight vs. in the plane of the sky, respectively. We find that the optical richness of round-BCG clusters is $\sim 10$\% larger than that of elliptical-BCG clusters, in agreement with the expectation from projection effects and presenting the first such detection in data. The density profiles, however, are not in agreement with the expectation from projection effects: the 1-halo term (below $6~h^{-1}\rm{Mpc}$) of both the weak-lensing and galaxy density profiles are the same for the subsamples, contrary to previous studies based on X-ray selected clusters. In the 2-halo regime (above $6~h^{-1}\rm{Mpc}$), we find a significant excess of the elliptical-BCG cluster profiles compared to the round-BCG cluster profiles, which is the opposite of the expectation from numerical simulations. We hypothesize that the intrinsic shape of the BCG reflects not just the orientation angle, but also intrinsic properties of the cluster which can affect both the SZ signal and the amplitude of the 2-halo term.

astro-ph.CO

Confronting cosmic shear astrophysical uncertainties: DES Year 3 revisited

Cosmology from weak gravitational lensing has been limited by astrophysical uncertainties in baryonic feedback and intrinsic alignments. By calibrating these effects using external data, we recover non-linear information, achieving a 2% constraint on the clustering amplitude, $S_8$, resulting in a factor of two improvement on the $\Lambda$CDM constraints relative to the fiducial Dark Energy Survey Year 3 model. The posterior, $S_8=0.832^{+0.013}_{-0.017}$, shifts by $1.5\sigma$ to higher values, in closer agreement with the cosmic microwave background result for the standard six-parameter $\Lambda$CDM cosmology. Our approach uses a star-forming 'blue' galaxy sample with intrinsic alignment model parameters calibrated by direct spectroscopic measurements, together with a baryonic feedback model informed by observations of X-ray gas fractions and kinematic Sunyaev-Zel'dovich effect profiles that span a wide range in halo mass and redshift. Our results provide a blueprint for next-generation surveys: leveraging galaxy properties to control intrinsic alignments and external gas probes to calibrate feedback, unlocking a substantial improvement in the precision of weak lensing surveys.

astro-ph.CO

T-FIX: Text-Based Explanations with Features Interpretable to eXperts

As LLMs are deployed in knowledge-intensive settings (e.g., surgery, astronomy, therapy), users are often domain experts who expect not just answers, but explanations that mirror professional reasoning. Yet evaluating whether an LLM "thinks like an expert" remains difficult: existing approaches rely on per-example expert annotation, making them costly, hard to scale, and tied to a single notion of correct reasoning within each domain. To address this gap, we introduce T-FIX, a unified evaluation framework that operationalizes expert alignment as a desired attribute of LLM-generated explanations. T-FIX spans seven scientific tasks across three domains, with each task evaluated against expert-defined criteria that capture domain-grounded reasoning rather than generic explanation quality. Our framework enables automatic, personalizable evaluation of expert alignment that generalizes to unseen explanations without ongoing expert involvement. Code is available at https://github.com/BrachioLab/FIX-2/.

cs.CL

Constraining Power of Wavelet vs. Power Spectrum Statistics for CMB Lensing and Weak Lensing with Learned Binning

We present forecasts for constraints on the matter density ($\Omega_m$) and the amplitude of matter density fluctuations at 8h$^{-1}$Mpc ($\sigma_8$) from CMB lensing convergence maps and galaxy weak lensing convergence maps. For CMB lensing convergence auto statistics, we compare the angular power spectra ($C_\ell$'s) to the wavelet scattering transform (WST) coefficients. For CMB lensing convergence $\times$ galaxy weak lensing convergence statistics, we compare the cross angular power spectra to wavelet phase harmonics (WPH). This work also serves as the first application of WST and WPH to these probes. For CMB lensing convergence, we find that WST and $C_\ell$'s yield similar constraints in forecasts for all surveys considered in this work. When CMB lensing convergence is crossed with galaxy weak lensing convergence projected from $\textit{Euclid}$ Data Release 2 (DR2), we find that WPH outperforms cross-$C_\ell$'s by factors between $2.2$ and $3.4$ for individual parameter constraints. To compare these different summary statistics, we develop a novel learned binning approach. This method compresses summary statistics while maintaining interpretability. We find this leads to improved constraints compared to more naive binning schemes for our wavelet-based statistics, but not for $C_\ell$'s. By learning the binning and measuring constraints on distinct data sets, our method is robust to overfitting by construction.

astro-ph.CO

PARROT: An Open Multilingual Radiology Reports Dataset

Rationale and Objectives: To develop and validate PARROT (Polyglottal Annotated Radiology Reports for Open Testing), a large, multicentric, open-access dataset of fictional radiology reports spanning multiple languages for testing natural language processing applications in radiology. Materials and Methods: From May to September 2024, radiologists were invited to contribute fictional radiology reports following their standard reporting practices. Contributors provided at least 20 reports with associated metadata including anatomical region, imaging modality, clinical context, and for non-English reports, English translations. All reports were assigned ICD-10 codes. A human vs. AI report differentiation study was conducted with 154 participants (radiologists, healthcare professionals, and non-healthcare professionals) assessing whether reports were human-authored or AI-generated. Results: The dataset comprises 2,658 radiology reports from 76 authors across 21 countries and 13 languages. Reports cover multiple imaging modalities (CT: 36.1%, MRI: 22.8%, radiography: 19.0%, ultrasound: 16.8%) and anatomical regions, with chest (19.9%), abdomen (18.6%), head (17.3%), and pelvis (14.1%) being most prevalent. In the differentiation study, participants achieved 53.9% accuracy (95% CI: 50.7%-57.1%) in distinguishing between human and AI-generated reports, with radiologists performing significantly better (56.9%, 95% CI: 53.3%-60.6%, p<0.05) than other groups. Conclusion: PARROT represents the largest open multilingual radiology report dataset, enabling development and validation of natural language processing applications across linguistic, geographic, and clinical boundaries without privacy constraints.

cs.CL

Field-Level Comparison and Robustness Analysis of Cosmological N-body Simulations

We present the first field-level comparison of cosmological N-body simulations, considering various widely used codes: Abacus, CUBEP$^3$M, Enzo, Gadget, Gizmo, PKDGrav, and Ramses. Unlike previous comparisons focused on summary statistics, we conduct a comprehensive field-level analysis: evaluating statistical similarity, quantifying implications for cosmological parameter inference, and identifying the regimes in which simulations are consistent. We begin with a traditional comparison using the power spectrum, cross-correlation coefficient, and visual inspection of the matter field. We follow this with a statistical out-of-distribution (OOD) analysis to quantify distributional differences between simulations, revealing insights not captured by the traditional metrics. We then perform field-level simulation-based inference (SBI) using convolutional neural networks (CNNs), training on one simulation and testing on others, including a full hydrodynamic simulation for comparison. We identify several causes of OOD behavior and biased inference, finding that resolution effects, such as those arising from adaptive mesh refinement (AMR), have a significant impact. Models trained on non-AMR simulations fail catastrophically when evaluated on AMR simulations, introducing larger biases than those from hydrodynamic effects. Differences in resolution, even when using the same N-body code, likewise lead to biased inference. We attribute these failures to a CNN's sensitivity to small-scale fluctuations, particularly in voids and filaments, and demonstrate that appropriate smoothing brings the simulations into statistical agreement. Our findings motivate the need for careful data filtering and the use of field-level OOD metrics, such as PQMass, to ensure robust inference.

astro-ph.CO

Map-level baryonification: unified treatment of weak lensing two-point and higher-order statistics

Precision cosmology benefits from extracting maximal information from cosmic structures, motivating the use of higher-order statistics (HOS) at small spatial scales. However, predicting how baryonic processes modify matter statistics at these scales has been challenging. The baryonic correction model (BCM) addresses this by modifying dark-matter-only simulations to mimic baryonic effects, providing a flexible, simulation-based framework for predicting both two-point and HOS. We show that a 3-parameter version of the BCM can jointly fit weak lensing maps' two-point statistics, wavelet phase harmonics coefficients, scattering coefficients, and the third and fourth moments to within 2% accuracy across all scales $\ell < 2000$ and tomographic bins for a DES-Y3-like redshift distribution ($z \lesssim 2$), using the FLAMINGO simulations. These results demonstrate the viability of BCM-assisted, simulation-based weak lensing inference of two-point and HOS, paving the way for robust cosmological constraints that fully exploit non-Gaussian information on small spatial scales.

astro-ph.CO

Reanalysis of Stage-III cosmic shear surveys: A comprehensive study of shear diagnostic tests

In recent years, shear catalogs have been released by various Stage-III weak lensing surveys including the Kilo-Degree Survey, the Dark Energy Survey, and the Hyper Suprime-Cam Subaru Strategic Program. These shear catalogs have undergone rigorous validation tests to ensure that the residual shear systematic effects in the catalogs are subdominant relative to the statistical uncertainties, such that the resulting cosmological constraints are unbiased. While there exists a generic set of tests that are designed to probe certain systematic effects, the implementations differ slightly across the individual surveys, making it difficult to make direct comparisons. In this paper, we use the TXPipe package to conduct a series of predefined diagnostic tests across three public shear catalogs -- the 1,000 deg$^2$ KiDS-1000 shear catalog, the Year 3 DES-Y3 shear catalog, and the Year 3 HSC-Y3 shear catalog. We attempt to reproduce the published results when possible and perform key tests uniformly across the surveys. While all surveys pass most of the null tests in this study, we find two tests where some of the surveys fail. Namely, we find that when measuring the tangential ellipticity around bright and faint star samples, KiDS-1000 fails with a $\chi^2$/dof of 121.1/16 and 257.7/16 for bins 4 and 5 for faint, weighted stars. We also find that DES-Y3 and HSC-Y3 fail the $B$-mode test when estimated with the Hybrid-$E$/$B$ method, with a $\chi^2$/dof of 37.9/10 and 36.0/8 for the fourth and third autocorrelation bins. We assess the impacts on the $\Omega_{\rm m}$ - S$_{8}$ parameter space by comparing the posteriors of a simulated data vector with and without PSF contamination -- we find negligible effects in all cases. Finally, we propose strategies for performing these tests on future surveys such as the Vera C. Rubin Observatory's Legacy Survey of Space and Time.

astro-ph.CO

Relationship between 2D and 3D Galaxy Stellar Mass and Correlations with Halo Mass

Recent studies suggest that the stars in the outer regions of massive galaxies trace halo mass better than the inner regions and that an annular stellar mass provides a low scatter method of selecting galaxy clusters. However, we can only observe galaxies as projected two-dimensional objects on the sky. In this paper, we use a sample of simulated galaxies to study how well galaxy stellar mass profiles in three dimensions correlate with halo mass, and what effects arise when observationally projecting stellar profiles into two dimensions. We compare 2D and 3D outer stellar mass selections and find that they have similar performance as halo mass proxies and that, surprisingly, a 2D selection sometimes has marginally better performance. We also investigate whether the weak lensing profiles around galaxies selected by 2D outer stellar mass suffer from projection effects. We find that the lensing profiles of samples selected by 2D and 3D definitions are nearly identical, suggesting that the 2D selection does not create a bias. These findings underscore the promise of using outer stellar mass as a tool for identifying galaxy clusters.

astro-ph.CO

The FIX Benchmark: Extracting Features Interpretable to eXperts

Feature-based methods are commonly used to explain model predictions, but these methods often implicitly assume that interpretable features are readily available. However, this is often not the case for high-dimensional data, and it can be hard even for domain experts to mathematically specify which features are important. Can we instead automatically extract collections or groups of features that are aligned with expert knowledge? To address this gap, we present FIX (Features Interpretable to eXperts), a benchmark for measuring how well a collection of features aligns with expert knowledge. In collaboration with domain experts, we propose FIXScore, a unified expert alignment measure applicable to diverse real-world settings across cosmology, psychology, and medicine domains in vision, language, and time series data modalities. With FIXScore, we find that popular feature-based explanation methods have poor alignment with expert-specified knowledge, highlighting the need for new methods that can better identify features interpretable to experts.

cs.LG

Dimensionality Reduction Techniques for Statistical Inference in Cosmology

We explore linear and non-linear dimensionality reduction techniques for statistical inference of parameters in cosmology. Given the importance of compressing the increasingly complex data vectors used in cosmology, we address questions that impact the constraining power achieved, such as: Are currently used methods effectively lossless? Under what conditions do nonlinear methods, typically based on neural nets, outperform linear methods? Through theoretical analysis and experiments with simulated weak lensing data vectors we compare three standard linear methods and neural network based methods. We propose two linear methods that outperform all others while using less computational resources: a variation of the MOPED algorithm we call e-MOPED and an adaptation of Canonical Correlation Analysis (CCA), which is a method new to cosmology but well known in statistics. Both e-MOPED and CCA utilize simulations spanning the full parameter space, and rely on the sensitivity of the data vector to the parameters of interest. The gains we obtain are significant compared to compression methods used in the literature: up to 30% in the Figure of Merit for $\Omega_m$ and $S_8$ in a realistic Simulation Based Inference analysis that includes statistical and systematic errors. We also recommend two modifications that improve the performance of all methods: First, include components in the compressed data vector that may not target the key parameters but still enhance the constraints on due to their correlations. The gain is significant, above 20% in the Figure of Merit. Second, compress Gaussian and non-Gaussian statistics separately -- we include two summary statistics of each type in our analysis.

astro-ph.CO

The Gravitational Lensing Imprints of DES Y3 Superstructures on the CMB: A Matched Filtering Approach

$ $Low density cosmic voids gravitationally lens the cosmic microwave background (CMB), leaving a negative imprint on the CMB convergence $\kappa$. This effect provides insight into the distribution of matter within voids, and can also be used to study the growth of structure. We measure this lensing imprint by cross-correlating the Planck CMB lensing convergence map with voids identified in the Dark Energy Survey Year 3 data set, covering approximately 4,200 deg$^2$ of the sky. We use two distinct void-finding algorithms: a 2D void-finder which operates on the projected galaxy density field in thin redshift shells, and a new code, Voxel, which operates on the full 3D map of galaxy positions. We employ an optimal matched filtering method for cross-correlation, using the MICE N-body simulation both to establish the template for the matched filter and to calibrate detection significances. Using the DES Y3 photometric luminous red galaxy sample, we measure $A_\kappa$, the amplitude of the observed lensing signal relative to the simulation template, obtaining $A_\kappa = 1.03 \pm 0.22$ ($4.6\sigma$ significance) for Voxel and $A_\kappa = 1.02 \pm 0.17$ ($5.9\sigma$ significance) for 2D voids, both consistent with $\Lambda$CDM expectations. We additionally invert the 2D void-finding process to identify superclusters in the projected density field, for which we measure $A_\kappa = 0.87 \pm 0.15$ ($5.9\sigma$ significance). The leading source of noise in our measurements is Planck noise, implying that future data from the Atacama Cosmology Telescope (ACT), South Pole Telescope (SPT) and CMB-S4 will increase sensitivity and allow for more precise measurements.

astro-ph.CO

Primordial non-Gaussianities with weak lensing: Information on non-linear scales in the Ulagam full-sky simulations

Primordial non-Gaussianities (PNGs) are signatures in the density field that encode particle physics processes from the inflationary epoch. Such signatures have been extensively studied using the Cosmic Microwave Background, through constraining the amplitudes, $f^{X}_{\rm NL}$, with future improvements expected from large-scale structure surveys; specifically, the galaxy correlation functions. We show that weak lensing fields can be used to achieve competitive and complementary constraints. This is shown via the new Ulagam suite of N-body simulations, a subset of which evolves primordial fields with four types of PNGs. We create full-sky lensing maps and estimate the Fisher information from three summary statistics measured on the maps: the moments, the cumulative distribution function, and the 3-point correlation function. We find that the year 10 sample from the Rubin Observatory Legacy Survey of Space and Time (LSST) can constrain PNGs to $σ(f^{\rm\,eq}_{\rm NL}) \approx 110$, $σ(f^{\rm\,or,lss}_{\rm NL}) \approx 120$, $σ(f^{\rm\,loc}_{\rm NL}) \approx 40$. For the former two, this is better than or comparable to expected galaxy clustering-based constraints from the Dark Energy Spectroscopic Instrument (DESI). The PNG information in lensing fields is on non-linear scales and at low redshifts ($z \lesssim 1.25$), with a clear origin in the evolution history of massive halos. The constraining power degrades by $\sim\!\!60\%$ under scale cuts of $\gtrsim 20{\,\rm Mpc}$, showing there is still significant information on scales mostly insensitive to small-scale systematic effects (e.g. baryons). We publicly release the Ulagam suite to enable more survey-focused analyses.

astro-ph.CO

Late Time Modification of Structure Growth and the S8 Tension

The $S_8$ tension between low-redshift galaxy surveys and the primary CMB signals a possible breakdown of the $Λ$CDM model. Recently differing results have been obtained using low-redshift galaxy surveys and the higher redshifts probed by CMB lensing, motivating a possible time-dependent modification to the growth of structure. We investigate a simple phenomenological model in which the growth of structure deviates from the $Λ$CDM prediction at late times, in particular as a simple function of the dark energy density. Fitting to galaxy lensing, CMB lensing, BAO, and Supernovae datasets, we find significant evidence - 2.5 - 3$σ$, depending on analysis choices - for a non-zero value of the parameter quantifying a deviation from $Λ$CDM. The preferred model, which has a slower growth of structure below $z\sim 1$, improves the joint fit to the data over $Λ$CDM. While the overall fit is improved, there is weak evidence for galaxy and CMB lensing favoring different changes in the growth of structure.

astro-ph.CO

Improving Convolutional Neural Networks for Cosmological Fields with Random Permutation

Convolutional Neural Networks (CNNs) have recently been applied to cosmological fields -- weak lensing mass maps and galaxy maps. However, cosmological maps differ in several ways from the vast majority of images that CNNs have been tested on: they are stochastic, typically low signal-to-noise per pixel, and with correlations on all scales. Further, the cosmology goal is a regression problem aimed at inferring posteriors on parameters that must be unbiased. We explore simple CNN architectures and present a novel approach of regularization and data augmentation to improve its performance for lensing mass maps. We find robust improvement by using a mixture of pooling and shuffling of the pixels in the deep layers. The random permutation regularizes the network in the low signal-to-noise regime and effectively augments the existing data. We use simulation-based inference (SBI) to show that the model outperforms CNN designs in the literature. We find a 30% improvement in the constraints of the $S_8$ parameter for simulated Stage-III surveys, including systematic uncertainties such as intrinsic alignments. We explore various statistical errors corresponding to next-generation surveys and find comparable improvements. We expect that our approach will have applications to other cosmological fields as well, such as galaxy maps or 21-cm maps.

astro-ph.CO

Sum-of-Parts: Self-Attributing Neural Networks with End-to-End Learning of Feature Groups

Self-attributing neural networks (SANNs) present a potential path towards interpretable models for high-dimensional problems, but often face significant trade-offs in performance. In this work, we formally prove a lower bound on errors of per-feature SANNs, whereas group-based SANNs can achieve zero error and thus high performance. Motivated by these insights, we propose Sum-of-Parts (SOP), a framework that transforms any differentiable model into a group-based SANN, where feature groups are learned end-to-end without group supervision. SOP achieves state-of-the-art performance for SANNs on vision and language tasks, and we validate that the groups are interpretable on a range of quantitative and semantic metrics. We further validate the utility of SOP explanations in model debugging and cosmological scientific discovery. Our code is available at https://github.com/BrachioLab/sop

cs.LG

Robust field-level inference with dark matter halos

We train graph neural networks on halo catalogues from Gadget N-body simulations to perform field-level likelihood-free inference of cosmological parameters. The catalogues contain $\lesssim$5,000 halos with masses $\gtrsim 10^{10}~h^{-1}M_\odot$ in a periodic volume of $(25~h^{-1}{\rm Mpc})^3$; every halo in the catalogue is characterized by several properties such as position, mass, velocity, concentration, and maximum circular velocity. Our models, built to be permutationally, translationally, and rotationally invariant, do not impose a minimum scale on which to extract information and are able to infer the values of $Ω_{\rm m}$ and $σ_8$ with a mean relative error of $\sim6\%$, when using positions plus velocities and positions plus masses, respectively. More importantly, we find that our models are very robust: they can infer the value of $Ω_{\rm m}$ and $σ_8$ when tested using halo catalogues from thousands of N-body simulations run with five different N-body codes: Abacus, CUBEP$^3$M, Enzo, PKDGrav3, and Ramses. Surprisingly, the model trained to infer $Ω_{\rm m}$ also works when tested on thousands of state-of-the-art CAMELS hydrodynamic simulations run with four different codes and subgrid physics implementations. Using halo properties such as concentration and maximum circular velocity allow our models to extract more information, at the expense of breaking the robustness of the models. This may happen because the different N-body codes are not converged on the relevant scales corresponding to these parameters.

astro-ph.CO