SearcharxivSearch

arXiv subjects

Leander Thiele

Publications and source records attributed to Leander Thiele.

At least 19 recordsLinked to original sources

Future of Artificial Intelligence for Science in Japan 2024 Community Report

This white paper summarizes scientific challenges and AI/ML research opportunities identified through the FAIRS Japan 2024 unconference process. The discussion focuses on three major physics domains: accelerator physics, cosmology and astrophysics, and neutrino physics. Although each domain has distinct scientific goals and experimental constraints, several common technical themes emerge: high-dimensional reconstruction, fast and accurate simulation, uncertainty propagation, simulation-to-data mismatch, anomaly detection, real-time decision-making, and shared infrastructure.

hep-ph

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.

astro-ph.IM

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.

cs.CL

Cluster Mass Inference from Galaxy Kinematics

The masses of galaxy clusters carry cosmological and astrophysical information. We develop a simulation-based inference pipeline to infer cluster masses from full projected phase-space information of member and interloper galaxies. Our method combines a permutation-invariant Deep Sets architecture with neural posterior estimation using normalizing flows, enabling the recovery of expressive posterior distributions. We train the model to predict residual corrections to the classical $M$--$\sigma$ relation, thus explicitly isolating information beyond velocity dispersion. Using the Uchuu-UniverseMachine simulation, we evaluate the method under both idealized (interloper-free) and realistic (cylindrical) observational setups. In the idealized case, our model reduces the scatter in mass estimates to as low as $\sim 0.1$ dex, representing a twofold improvement over the traditional $M$--$\sigma$ relation. In the cylindrical setup, we achieve comparable performance at the high-mass end ($> 10^{14.5}\,M_\odot/h$), demonstrating robustness against interloper contamination. We demonstrate that set-based simulation-driven inference provides a powerful and flexible framework for galaxy cluster mass estimation, enabling improved accuracy and reliable uncertainty characterization for upcoming large-scale surveys. Our model saturates the kinematic information content and thus suggests a baseline for future studies.

astro-ph.CO

Machine Learning Techniques for Astrophysics and Cosmology: Simulation-Based Inference

Simulation-based inference (SBI) enables parameter inference by training neural networks on forward simulations. It is being applied both for intractable likelihoods as well as under time constraints on the posterior sampling. After motivating situations in which SBI is useful, we give a pedagogical description of the basic techniques. These are posterior, likelihood, and ratio estimation. Alternatives, sequential versions, and learned summaries are discussed briefly. We provide a brief guide to choosing among the techniques in practical scenarios. SBI needs to be verified through diagnostics since failures can be subtle but would invalidate the inference result. We explain the most common diagnostic techniques. We briefly list some recent SBI applications in the cosmology and astrophysics literature. Before concluding, we discuss current methodological challenges. We identify training with limited simulation budgets as the critical problem for applications to cosmology and astrophysics.

astro-ph.CO

LLMs with in-context learning for Algorithmic Theoretical Physics

There is an increasing number of algorithmic computations in theoretical physics. These, while conceptually simple, can nevertheless be time-consuming and contain subtleties that should not be overlooked. Given the recent improvement of Large Language Models (LLM), it is natural to investigate whether LLMs equipped with a computer algebra system (CAS) runtime and sufficiently informative context can reliably carry out these algorithmic tasks. In this work, we interface Claude with Maple, and apply this framework to cosmological perturbations in modified theories of gravity. We demonstrate the current capabilities of this approach, the typical failures, and how the same can be improved. We find that a frontier LLM supplied with worked examples is able to solve most test problems.

cs.LG

Bayesian Cosmic Void Finding with Graph Flows

Cosmic voids contain higher-order cosmological information and are of interest for astroparticle physics. Finding genuine matter underdensities in sparse galaxy surveys is, however, an underconstrained problem. Traditional void finding algorithms produce deterministic void catalogs, neglecting the probabilistic nature of the problem. We present a method to sample from the stochastic mapping from galaxy catalogs to arbitrary void definitions. Our algorithm uses a deep graph neural network to evolve "test particles" according to a flow-matching objective. We demonstrate the method in a simplified example setting but outline steps to generalize it towards practically usable void finders. Trained on a deterministic teacher, the model performs well but has considerable stochasticity which we interpret as regularization. Cosmological information in the predicted void catalogs outperforms the teacher. On the one hand, our method can cheaply emulate existing void finders with apparently useful regularization. More importantly, it also allows us to find the Bayes-optimal mapping between observed galaxies and any void definition. This includes definitions operating at the level of simulated matter density and velocity fields.

astro-ph.CO

Replicating weak-lensing summary-statistic covariances with normalizing flows

We explore the ability of normalizing flow (NF) generative models to reproduce weak-lensing summary statistics when trained on a set of cosmological simulations. Our analysis focuses on how accurately NF models recover the mean, standard deviation, and covariance of key statistics derived from convergence ($\kappa$) maps: The angular power spectrum $C_{\ell}$, probability density function, and Minkowski functionals of weak lensing convergence $\kappa$-maps. We test two scenarios for training: (1) on the data vectors and (2) on the full $\kappa$-maps. In both cases, the NF models reproduce the mean and variance of the target statistics within percent-level accuracy. However, the accuracy of the off-diagonal elements of the covariance matrix is underestimated by up to $\sim25\%$. We study several mitigation strategies and find that data augmentation and training with noisy fields help improve covariance recovery to $\mathcal{O}(5\%)$ on power spectrum statistics. Our study demonstrates that while the means and variances of weak lensing statistics can be well modeled by NF, covariances can be significantly underestimated if mitigation strategies are not applied. We present this test as a rigorous diagnostic of generative-model fidelity.

astro-ph.CO

A new constraint on the $y$-distortion with FIRAS: implications for feedback models in galaxy formation and cosmic shear measurements

The $y$-type distortion of the blackbody spectrum of the cosmic microwave background radiation probes the pressure of the gas trapped in galaxy groups and clusters. We reanalyze archival data of the FIRAS instrument with an improved astrophysical foreground cleaning technique, and measure a mean $y$-distortion of $\langle y\rangle = (1.2\pm 2.0) \times 10^{-6}$ ($\langle y\rangle\lesssim 5.2\times 10^{-6}$ at 95\% C.L.), a factor of $\sim 3$ tighter than the original FIRAS results. This measurement directly rules out many models of baryonic feedback as implemented in cosmological hydrodynamical simulations, mostly using information in objects with mass $M\lesssim 10^{14} {\rm M}_{\odot}$. We discuss its implications for the analysis of cosmic shear and kinetic Sunyaev-Zel'dovich effect data, and future spectral distortion experiments.

astro-ph.CO

Reconstructing the local density field with combined convolutional and point cloud architecture

We construct a neural network to perform regression on the local dark-matter density field given line-of-sight peculiar velocities of dark-matter halos, biased tracers of the dark matter field. Our architecture combines a convolutional U-Net with a point-cloud DeepSets. This combination enables efficient use of small-scale information and improves reconstruction quality relative to a U-Net-only approach. Specifically, our hybrid network recovers both clustering amplitudes and phases better than the U-Net on small scales.

astro-ph.CO

Impact of Simulation Box Size for Weak Lensing: Replication and Super-Sample Effects

We quantify the bias caused by small simulation box size on weak lensing observables and covariances, considering both replication and super-sample effects for a range of higher-order statistics. Using two simulation suites -- one comprising large boxes ($3750\,h^{-1}{\rm Mpc}$) and another constructed by tiling small boxes ($625\,h^{-1}{\rm Mpc}$) -- we generate full-sky convergence maps and extract $10^\circ \times 10^\circ$ patches via a Fibonacci grid. We consider biases in the mean and covariance of the angular power spectrum, bispectrum (up to $\ell=3000$), PDF, peak/minima counts, and Minkowski functionals. By first identifying lines of sight that are impacted by replications, we find that replication causes a O$(10\%)$ bias in the PDF and Minkowski functionals, and a O$(1\%)$ bias in other summary statistics. Replication also causes a O$(10\%)$ bias in the covariances, increasing with source redshift and $\ell$, reaching $\sim25\%$ for $z_s=2.5$. We additionally show that replication leads to heavy biases (up to O$(100\%)$ at high redshift) when performing gnomonic projection on a patch that is centered along a direction of replication. We then identify the lines of sight that are minimally affected by replication, and use the corresponding patches to isolate and study super-sample effects, finding that, while the mean values agree to within $1\%$, the variances differ by O$(10\%)$ for $z_s\leq2.5$. We show that these effects remain in the presence of noise and smoothing scales typical of the DES, KiDS, HSC, LSST, Euclid, and Roman surveys. We also discuss how these effects scale as a function of box size. Our results highlight the importance of large simulation volumes for accurate lensing statistics and covariance estimation.

astro-ph.CO

Set-based Implicit Likelihood Inference of Galaxy Cluster Mass

We present a set-based machine learning framework that infers posterior distributions of galaxy cluster masses from projected galaxy dynamics. Our model combines Deep Sets and conditional normalizing flows to incorporate both positional and velocity information of member galaxies to predict residual corrections to the $M$-$σ$ relation for improved interpretability. Trained on the Uchuu-UniverseMachine simulation, our approach significantly reduces scatter and provides well-calibrated uncertainties across the full mass range compared to traditional dynamical estimates.

cs.LG

First Constraints from Marked Angular Power Spectra with Subaru Hyper Suprime-Cam Survey First-Year Data

We present the first application of marked angular power spectra to weak lensing data, using maps from the Subaru Hyper Suprime-Cam Year 1 (HSC-Y1) survey. Marked convergence fields, constructed by weighting the convergence field with non-linear functions of its smoothed version, are designed to encode higher-order information while remaining computationally tractable. Using simulations tailored to the HSC-Y1 data, we test three mark functions that up- or down-weight different density environments. Our results show that combining multiple types of marked auto- and cross-spectra improves constraints on the clustering amplitude parameter $S_8\equivσ_8\sqrt{Ω_{\rm m}/0.3}$ by $\approx$43\% compared to standard two-point power spectra. When applied to the HSC-Y1 data, this translates into a constraint on $S_8 = 0.807\pm 0.024$. We assess the sensitivity of the marked power spectra to systematics, including baryonic effects, intrinsic alignment, photometric redshifts, and multiplicative shear bias. These results demonstrate the promise of marked statistics as a practical and powerful tool for extracting non-Gaussian information from weak lensing surveys.

astro-ph.CO

Simulation-Efficient Cosmological Inference with Multi-Fidelity SBI

The simulation cost for cosmological simulation-based inference can be decreased by combining simulation sets of varying fidelity. We propose an approach to such multi-fidelity inference based on feature matching and knowledge distillation. Our method results in improved posterior quality, particularly for small simulation budgets and difficult inference problems.

astro-ph.CO

The Primordial Inflation Explorer (PIXIE): Mission Design and Science Goals

The Primordial Inflation Explorer (PIXIE) is an Explorer-class mission concept to measure the energy spectrum and linear polarization of the cosmic microwave background (CMB). A single cryogenic Fourier transform spectrometer compares the sky to an external blackbody calibration target, measuring the Stokes I, Q, U parameters to levels ~200 Jy/sr in each 2.65 degree diameter beam over the full sky, in each of 300 frequency channels from 28 GHz to 6 THz. With sensitivity over 1000 times greater than COBE/FIRAS, PIXIE opens a broad discovery space for the origin, contents, and evolution of the universe. Measurements of small distortions from a CMB blackbody spectrum provide a robust determination of the mean electron pressure and temperature in the universe while constraining processes including dissipation of primordial density perturbations, black holes, and the decay or annihilation of dark matter. Full-sky maps of linear polarization measure the optical depth to reionization at nearly the cosmic variance limit and constrain models of primordial inflation. Spectra with sub-percent absolute calibration spanning microwave to far-IR wavelengths provide a legacy data set for analyses including line intensity mapping of extragalactic emission and the cosmic infrared background amplitude and anisotropy. We describe the PIXIE instrument sensitivity, foreground subtraction, and anticipated science return from both the baseline 2-year mission and a potential extended mission.

astro-ph.CO

Cosmological constraints using Minkowski functionals from the first year data of the Hyper Suprime-Cam

We use Minkowski functionals to analyse weak lensing convergence maps from the first-year data release of the Subaru Hyper Suprime-Cam (HSC-Y1) survey. Minkowski functionals provide a description of the morphological properties of a field, capturing the non-Gaussian features of the Universe matter-density distribution. Using simulated catalogs that reproduce survey conditions and encode cosmological information, we emulate Minkowski functionals predictions across a range of cosmological parameters to derive the best-fit from the data. By applying multiple scales cuts, we rigorously mitigate systematic effects, including baryonic feedback and intrinsic alignments. From the analysis, combining constraints of the angular power spectrum and Minkowski functionals, we obtain $S_8 \equiv σ_8\sqrt{Ω_{\rm m}/0.3} = {0.808}_{-0.046}^{+0.033}$ and $Ω_{\rm m} = {0.293}_{-0.043}^{+0.157}$. These results represent a $40\%$ improvement on the $S_8$ constraints compared to using power spectrum only. \newtext{Minkowski functionals results are consistent with other two-point, and higher order statistics constraints using the same data, being in agreement with CMB results from the Planck $S_8$ measurements. Our study demonstrates the power of Minkowski functionals beyond two-point statistics to constrain and break the degeneracy between $Ω_{\rm m}$ and $σ_8$.

astro-ph.CO

De-baryonifying halos via optimal transport

Baryonic feedback uncertainty is a limiting systematic for next-generation weak gravitational lensing analyses. At the same time, high-resolution weak lensing maps are best analyzed at the field-level. Thus, robustly accounting for the baryonic effects in the projected matter density field is required. Ideally, constraints on feedback strength from astrophysical probes should be folded into the weak lensing field-level likelihood. We propose a macroscopic method based on an empirical correlation between feedback strength and an optimal transport cost. Since feedback is local re-distribution of matter, optimal transport is a promising concept. In this proof-of-concept, we de-baryonify projected mass around individual halos in the IllustrisTNG simulation. We choose the de-baryonified solution as the point of maximum likelihood on the hypersurface defined by fixed optimal transport cost around the observed full-physics halos. The likelihood is approximated through a normalizing flow trained on multiple gravity-only simulations. We find that the set of de-baryonified halos reproduces the correct convergence power spectrum suppression. There is considerable scatter when considering individual halos. We outline how the optimal transport de-baryonification concept can be generalized to full convergence maps.

astro-ph.CO

Cosmology from HSC Y1 Weak Lensing with Combined Higher-Order Statistics and Simulation-based Inference

We present cosmological constraints from weak lensing with the Subaru Hyper Suprime-Cam (HSC) first-year (Y1) data, using a simulation-based inference (SBI) method. % We explore the performance of a set of higher-order statistics (HOS) including the Minkowski functionals, counts of peaks and minima, and the probability distribution function and compare them to the traditional two-point statistics. The HOS, also known as non-Gaussian statistics, can extract additional non-Gaussian information that is inaccessible to the two-point statistics. We use a neural network to compress the summary statistics, followed by an SBI approach to infer the posterior distribution of the cosmological parameters. We apply cuts on angular scales and redshift bins to mitigate the impact of systematic effects. Combining two-point and non-Gaussian statistics, we obtain $S_8 \equiv σ_8 \sqrt{Ω_m/0.3} = 0.804_{-0.040}^{+0.041}$ and $Ω_m = 0.344_{-0.090}^{+0.083}$, similar to that from non-Gaussian statistics alone. These results are consistent with previous HSC analyses and Planck 2018 cosmology. Our constraints from non-Gaussian statistics are $\sim 25\%$ tighter in $S_8$ than two-point statistics, where the main improvement lies in $Ω_m$, with $\sim 40$\% tighter error bar compared to using the angular power spectrum alone ($S_8 = 0.766_{-0.056}^{+0.054}$ and $Ω_m = 0.365_{-0.141}^{+0.148}$). We find that, among the non-Gaussian statistics we studied, the Minkowski functionals are the primary driver for this improvement. Our analyses confirm the SBI as a powerful approach for cosmological constraints, avoiding any assumptions about the functional form of the data's likelihood.

astro-ph.CO