SearcharxivSearch

arXiv subjects

Caroline Heneka

Publications and source records attributed to Caroline Heneka.

At least 19 recordsLinked to original sources

Cross-simulator transfer with foundation model summaries: Towards robust SKA-era reionization inference

Simulation-based inference (SBI) for parameter estimation is vulnerable to model misspecification: neural summaries and density estimators trained on a specific forward model typically fail when applied to data drawn from another model, or from real observations, and no training simulator can capture the full observational pipeline of a real measurement exactly. We show that a self-supervised Vision Transformer (ViT), pretrained label-free on a fast approximate simulator, produces transferable data summaries that generalize across simulators. Without retraining, it can be reused as a frozen encoder to infer astrophysical parameters from a completely different simulator that resolves the radiative transfer explicitly, on which it has never seen either data or parameters. As a concrete use case in 21cm cosmology, SKATR, a ViT pretrained with a Joint Embedding Predictive Architecture (JEPA), serves as a foundation model for reionization inference from upcoming SKA measurements: SKATR is pretrained once on 67k low-cost, noiseless semi-numerical 21cmFAST lightcones, then frozen and applied to hydrodynamical Loreli II lightcones, where a lightweight conditional flow matching head infers five astrophysical parameters; the encoder is never shown Loreli data, its parameters, or any noise. In our comparison, SKATR yields the most precise and best-calibrated posteriors across all five parameters, matching the accuracy of the fully-supervised in-domain baseline while requiring 2.6x fewer radiative-transfer simulations. Under realistic SKA AA* noise, only SKATR remains simultaneously accurate, informative, and calibrated, outperforming even a supervised baseline retrained from scratch on noisy data. Self-supervised pretraining on computationally efficient semi-numerical simulations is therefore a viable route to calibrated, simulator- and noise-agnostic reionization inference for the SKA-era.

astro-ph.CO

Square Kilometer Array Synergies for the Epoch of Reionization and Cosmic Dawn

Synergies with other instruments will be essential in making, verifying, and interpreting a detection of the cosmic 21-cm signal from the Epoch of Reionization (EoR) and Cosmic Dawn (CD) with the Square Kilometer Array (SKA) telescope. Such synergies can (i) provide prior information about galaxies and the intergalactic medium (IGM) during the EoR/CD; (ii) pave the road to a first 21cm detection by mitigating foregrounds and systematics through cross-correlations; and (iii) give complimentary physical insights into the galaxy -- IGM connection. Here we review the current state of synergies and discuss what observations will best compliment SKA-low EoR/CD observations.

astro-ph.CO

Inferring Cosmology and Astrophysics from the High-redshift 21cm Signal with SKA-Low

The Square Kilometre Array's low frequency telescope (SKA-Low) will enable inference of astrophysical and cosmological parameters from the redshifted 21 cm signal, probing the Cosmic Dawn and Epoch of Reionisation. While the power spectrum is the primary target for initial detection, the inherently non-Gaussian nature of the 21 cm signal, driven by the patchy evolution of ionised regions and spin temperature fluctuations, encodes rich information accessible through higher-order statistics and morphological measurements. Extracting these constraints requires diverse inference tools, encompassing both sophisticated modelling frameworks (analytical, semi-numerical, numerical, and emulators) used to predict the 21 cm signal, and advanced inference techniques (Bayesian, simulation-based, field-level) to connect statistics to the underlying physics. This chapter reviews these tools and explores the constraining power of different statistical probes accessible with SKA-Low, including the power spectrum, statistics beyond order two, moments of the signal distribution, and morphological measures. Combining these complementary statistics is crucial for breaking parameter degeneracies and unveiling the properties of the early Universe. We specifically assess the potential of the initial SKA-Low configuration (AA*) to measure galaxy and IGM properties, demonstrating its capability for early science results. This chapter forms part of a comprehensive set detailing the Epoch of Reionisation and Cosmic Dawn science case for the SKA-Low telescope.

astro-ph.CO

Overview of 21cm Experiments at high redshift with SKAO

We provide an overview of the eight SKAO Science Book chapters that motivate the Epoch of Reionisation and Cosmic Dawn experiments with SKA-Low. We describe the individual SKA-Low experiments and expected sensitivity - power spectrum, tomography, 21-cm forest, cross-correlations, building on the broad observational plan laid out in the 2015 SKA Science Book. Finally, we outline features of the telescope that will be critical for the success of EoR/CD science, e.g., beam apodization, substations, and multi-beaming.

astro-ph.CO

Cosmology with galaxy clusters using machine learning. Application to eROSITA Data

Context: We present the first Cosmological Parameter inferences from eROSITA X-ray observations of galaxy clusters using a Machine Learning algorithm. Methods: We train a Random Forest using mock catalogs of clusters from Magneticum multi-cosmology hydrodynamical simulations. We apply the trained ML algorithm to observed X-ray features (gas luminosity, mass, and temperature) at different redshifts from the eROSITA eFEDS and eRASS1 catalogs. Results: We obtain cosmological constraints with precision comparable to those from standard analyses, such as weak lensing and cluster abundances. We infer $\Omega_{\rm m}=0.30^{+0.03}_{-0.02}$, $\sigma_8=0.81\pm0.01$, and $h_0=0.710\pm0.004$. The recovered parameters show no tension in the $\Omega_{\rm m}-\sigma_8$ space, but a significant deviation of $h_0$ from the Planck estimates. These inferences remain rather stable against variations of the input observable set and parameter space coverage. These results indicate that correlations among intracluster properties contain cosmological information beyond that encoded in the cluster abundance alone, which can be captured by machine learning trained on multi-cosmology simulations. Conclusions: ML algorithms trained on multi-cosmology hydrodynamical simulations can effectively infer cosmological parameters directly from galaxy cluster data. This is a change of paradigm in the context of cosmological parameter inferences. This approach complements traditional cluster-count analyses and is particularly suited to large upcoming surveys, where systematic uncertainties in mass calibration may otherwise dominate the error budget. It also highlights the potential of large-scale X-ray surveys to deliver independent tests of the standard cosmological model.

astro-ph.CO

Constraining reionization morphology and source properties with 21cm galaxy cross-correlation surveys

Cross-correlations between 21cm observations and galaxy surveys provide a powerful probe of reionization by providing robustness against foreground contamination while linking ionization morphology to galaxies. We quantified the constraining power of 21cm galaxy cross-power spectra for inferring the neutral hydrogen fraction, $x_\mathrm{HI}(z),$ and mean overdensity, $\langle 1+\delta_\mathrm{HI} \rangle(z)$, exploring dependence on the field of view; redshift precision, $\sigma_z$; and minimum halo mass, $M_\mathrm{h,min}$. We employed our simulation-based inference framework EoRFlow for likelihood-free parameter estimation. Mock observations include thermal noise for 100h of SKA-Low with foreground avoidance and realistic galaxy-survey effects. For a fiducial survey ($\mathrm{FOV}=100\,\mathrm{deg}^2$, $\sigma_z=0.001$, $M_\mathrm{h,min}=10^{11}\mathrm{M}_\odot$), cross-power spectra yield unbiased constraints with posterior volumes (PVs) of $\sim$10% relative to priors. Cross-power measurements reduce the PV by 20-30% versus 21cm auto-power alone. With foreground avoidance, spectroscopic redshift precision is essential; photometric redshifts render cross-correlations uninformative. Notably, cross-power spectra constrain ionizing source properties, the escape fraction $f_\mathrm{esc,}$ and the star formation efficiency $f_*$, which remain degenerate in auto-power (PV >60%). Tight constraints require either deep surveys detecting faint galaxies ($M_\mathrm{h,min} \sim 10^{10}\mathrm{M}_\odot$) with moderate foregrounds (PV~11%) or conservative mass limits with optimistic foreground removal (PV~19%). 21cm galaxy cross-correlations enhance morphology constraints beyond auto-power while enabling previously inaccessible source property constraints. Realizing full potential requires precise redshifts and either faint galaxy detection limits or improved 21cm foreground cleaning.

astro-ph.CO

The 21cm-galaxy cross-correlation: Realistic forecast for 21cm signal detection and reionisation constraints

21cm-galaxy cross-correlation will play a key role in confirming the cosmological 21cm signal. We investigate which survey configurations detect the 21cm-LAE cross-correlation signal, and assess its ability to distinguish reionisation scenarios. Our pipeline computes observational uncertainties for the 21cm-galaxy cross-power spectrum, accounting for key survey parameters: the field of view (FoV), limiting luminosity of galaxy surveys $L_\alpha$, redshift uncertainty $\sigma_z$, and 21cm foreground wedge assumptions. We calculate the signal-to-noise ratio (SNR) of the 21cm-Lyman-$\alpha$ emitter (LAE) cross-power spectrum for two scenarios: one where reionisation is driven by faint or by bright galaxies. We find: (i) SNR increases with larger FoV, fainter $L_\alpha$, and smaller $\sigma_z$, with the FoV having the strongest impact when $\sigma_z$ is small. (ii) Under a moderate foreground wedge, photometric-like surveys yield insufficient SNR, and medium-deep ($L_\alpha\gtrsim10^{42.5}$erg s$^{-1}$), wide-area (FoV>20deg$^2$) slitless spectroscopic surveys are needed. (iii) Under an optimistic foreground wedge, detection is possible with deep ($L_\alpha\gtrsim10^{42.3}$erg s$^{-1}$), wide-area (FoV$\gtrsim80$deg$^2$) photometric-like or shallower, small-area (FoV$\simeq2-3$deg$^2$) slitless spectroscopic surveys. (iv) To distinguish the two reionisation scenarios at z=7, moderate foreground wedge scenarios require deep-wide spectroscopic surveys; under an optimistic foreground wedge, shallower, medium-area (FoV$\simeq10$deg$^2$) slitless spectroscopic surveys suffice. (v) Maximising the SNR for detection and model discrimination requires sampling the large-scale peak of the cross-power spectrum, which shifts to larger physical scales as reionisation proceeds and the less ionisation fronts follow the gas density - making surveys at z>7 more promising despite lower galaxy number densities.

astro-ph.CO

Starobinsky in Stereo: SKA-CMB Synergy in SBI

Modern machine learning techniques can unlock the vast cosmological information encoded in forthcoming Square Kilometre Array (SKA) observations. We show that tomographic 21 cm data from the reionisation era can yield stringent tests of inflationary models - here illustrated with Starobinsky $R+R^2$ inflation. Using a simulation-based inference (SBI) framework, we compare neural summaries (convolutional network and vision transformer) with a traditional power spectrum summary and perform a fully joint SBI analysis combining 21 cm data with data of the cosmic microwave background (CMB). Forecasts based on realistic mock observations indicate that SKA alone will achieve constraints competitive with Planck, and that the combined SKA + CMB dataset will tighten bounds on both inflationary and $\Lambda\mathrm{CDM}$ parameters considerably while improving precision on key astrophysical quantities.

astro-ph.CO

Direct reconstruction of the Reionization history from 21cm 2D Power Spectra

The 21cm line from the spin-flip transition of neutral hydrogen (HI) provides a unique window into the Epoch of Reionization (EoR), the final phase transition of our Universe. The Square Kilometre Array (SKA) enables precise measurements of 21cm fluctuations that trace ionization, temperature, and density fluctuations of the intergalactic medium (IGM). Nevertheless, a direct reconstruction of the timeline of the EoR in terms of the progress of ionization remains an ongoing challenge due to the highly non-Gaussian nature and thus intractable likelihood of the 21cm signal. Here, we present EoRFlow, a simulation-based inference (SBI) framework for reconstructing the global neutral hydrogen fraction $x_{\mathrm{HI}}(z)$ directly from 2D cylindrically averaged power spectra (2DPS) of the 21cm signal. We validate our method on realistic mock datasets for SKA-Low. Bypassing the need for explicit likelihood formulations, our approach enables fast, unbiased posterior estimation of the $x_{\mathrm{HI}}$ evolution in narrow redshift slices, allowing for piecewise reconstruction of the global reionization history. By directly inferring the reionization history from 21cm power spectra, our framework provides a scalable and robust path forward for 21cm cosmology in the SKA era.

astro-ph.CO

Large Language Models -- the Future of Fundamental Physics?

For many fundamental physics applications, transformers, as the state of the art in learning complex correlations, benefit from pretraining on quasi-out-of-domain data. The obvious question is whether we can exploit Large Language Models, requiring proper out-of-domain transfer learning. We show how the Qwen2.5 LLM can be used to analyze and generate SKA data, specifically 3D maps of the cosmological large-scale structure for a large part of the observable Universe. We combine the LLM with connector networks and show, for cosmological parameter regression and lightcone generation, that this Lightcone LLM (L3M) with Qwen2.5 weights outperforms standard initialization and compares favorably with dedicated networks of matching size.

astro-ph.CO

Cosmology from LOFAR Two-metre Sky Survey Data Release 2: Cross-correlations with luminous red galaxies from eBOSS

We cross-correlated galaxies from the LOw-Frequency ARray (LOFAR) Two-metre Sky Survey (LoTSS) second data release (DR2) radio source with the extended Baryon Oscillation Spectroscopic Survey (eBOSS) luminous red galaxy (LRG) sample to extract the baryon acoustic oscillation (BAO) signal and constrain the linear clustering bias of radio sources in LoTSS DR2. In the LoTSS DR2 catalogue, employing a flux density limit of $1.5$ mJy at the central LoTSS frequency of 144 MHz and a signal-to-noise ratio (S/N) of $7.5$, additionally considering eBOSS LRGs with redshifts between 0.6 and 1, we measured both the angular LoTSS-eBOSS cross-power spectrum and the angular eBOSS auto-power spectrum. These measurements were performed across various eBOSS redshift tomographic bins with a width of $\Delta z=0.06$. By marginalising over the broadband shape of the angular power spectra, we searched for a BAO signal in cross-correlation with radio galaxies, and determine the linear clustering bias of LoTSS radio sources for a constant-bias and an evolving-bias model. Using the cross-correlation, we measured the isotropic BAO dilation parameter as $\alpha=1.01\pm 0.11$ at $z_{\rm eff}=0.63$. By combining four redshift slices at $z_{\rm eff}=0.63, 0.69, 0.75$, and $0.81$, we determined a more constrained value of $\alpha = 0.968^{+0.060}_{-0.095}$. For the entire redshift range of $z_{\rm eff}=0.715$, we measured $b_C = 2.64 \pm 0.20$ for the constant-bias model, $b(z)=b_C$, and then $b_D = 1.80 \pm 0.13$ for the evolving-bias model, $b(z) = b_D / D(z)$, with $D(z)$ denoting the growth rate of linear structures. Additionally, we measured the clustering bias for individual redshift bins.

astro-ph.CO

Cosmology from LOFAR Two-metre Sky Survey Data Release 2: Counts-in-Cells Statistics

We investigate the statistical distribution of source counts-in-cells in the second data release of the LOFAR Two-Metre Sky Survey (LoTSS-DR2) and we test a computationally cheap method based on the counts-in-cells to estimate the two-point correlation function. We compare three stochastic models for the counts-in-cells which result in a Poisson distribution, a compound Poisson distribution, and a negative binomial distribution. By analysing the variance of counts-in-cells for various cell sizes, we fit the reduced normalised variance to a single power-law model representing the angular two-point correlation function. Our analysis confirms that radio sources are not Poisson distributed, which is most likely due to multiple physical components of radio sources. Employing instead a Cox process, we show that there is strong evidence in favour of the negative binomial distribution above a flux density threshold of 2 mJy. Additionally, the mean number of radio components derived from the negative binomial distribution is in good agreement with corresponding estimates based on the value-added catalogue of LoTSS-DR2. The scaling of the counts-in-cells normalised variance with cell size is in good agreement with a power-law model for the angular two-point correlation. At a flux density threshold of 2 mJy and a signal-to-noise ratio of 7.5 for individual radio sources, we find that for a range of angular scales large enough to not be affected by the multi-component nature of radio sources, the value of the exponent of the power law ranges from -0.8 to -1.05. This closely aligns with findings from previous optical, infrared, and radio surveys of the large scale structure. The scaling of the counts-in-cells statistics with cell size provides a computationally efficient method to estimate the two-point correlation properties, offering a valuable tool for future large-scale structure studies.

astro-ph.CO

Galaxy Spectra Networks (GaSNet). III. Generative pre-trained network for spectrum reconstruction, redshift estimate and anomaly detection

Classification of spectra (1) and anomaly detection (2) are fundamental steps to guarantee the highest accuracy in redshift measurements (3) in modern all-sky spectroscopic surveys. We introduce a new Galaxy Spectra Neural Network (GaSNet-III) model that takes advantage of generative neural networks to perform these three tasks at once with very high efficiency. We use two different generative networks, an autoencoder-like network and U-Net, to reconstruct the rest-frame spectrum (after redshifting). The autoencoder-like network operates similarly to the classical PCA, learning templates (eigenspectra) from the training set and returning modeling parameters. The U-Net, in contrast, functions as an end-to-end model and shows an advantage in noise reduction. By reconstructing spectra, we can achieve classification, redshift estimation, and anomaly detection in the same framework. Each rest-frame reconstructed spectrum is extended to the UV and a small part of the infrared (covering the blueshift of stars). Owing to the high computational efficiency of deep learning, we scan the chi-squared value for the entire type and redshift space and find the best-fitting point. Our results show that generative networks can achieve accuracy comparable to the classical PCA methods in spectral modeling with higher efficiency, especially achieving an average of $>98\%$ classification across all classes ($>99.9\%$ for star), and $>99\%$ (stars), $>98\%$ (galaxies) and $>93\%$ (quasars) redshift accuracy under cosmology research requirements. By comparing different peaks of chi-squared curves, we define the ``robustness'' in the scanned space, offering a method to identify potential ``anomalous'' spectra. Our approach provides an accurate and high-efficiency spectrum modeling tool for handling the vast data volumes from future spectroscopic sky surveys.

astro-ph.GA

SKATR: A Self-Supervised Summary Transformer for SKA

The Square Kilometer Array will initiate a new era of radio astronomy by allowing 3D imaging of the Universe during Cosmic Dawn and Reionization. Modern machine learning is crucial to analyse the highly structured and complex signal. However, accurate training data is expensive to simulate, and supervised learning may not generalize. We introduce a self-supervised vision transformer, SKATR, whose learned encoding can be cheaply adapted for downstream tasks on 21cm maps. Focusing on regression and generative inference of astrophysical and cosmological parameters, we demonstrate that SKATR representations are maximally informative and that SKATR generalises out-of-domain to differently-simulated, noised, and higher-resolution datasets.

astro-ph.IM

Optimal, fast, and robust inference of reionization-era cosmology with the 21cmPIE-INN

Modern machine learning will allow for simulation-based inference from reionization-era 21cm observations at the Square Kilometre Array. Our framework combines a convolutional summary network and a conditional invertible network through a physics-inspired latent representation. It allows for an efficient and extremely fast determination of the posteriors of astrophysical and cosmological parameters, jointly with well-calibrated and on average unbiased summaries. The sensitivity to non-Gaussian information makes our method a promising alternative to the established power spectra.

astro-ph.CO

Deep Learning 21cm Lightcones in 3D

Interferometric measurements of the 21cm signal are a prime example of the data-driven era in astrophysics we are entering with current and upcoming experiments. We showcase the use of deep networks that are tailored for the structure of 3D tomographic 21cm light-cones to firstly detect and characterise HI sources and to secondly directly infer global astrophysical and cosmological model parameters. We compare different architectures and highlight how 3D CNN architectures that mirror the data structure are the best-performing model.

astro-ph.CO

Galaxy Spectra neural Network (GaSNet). II. Using Deep Learning for Spectral Classification and Redshift Predictions

Large sky spectroscopic surveys have reached the scale of photometric surveys in terms of sample sizes and data complexity. These huge datasets require efficient, accurate, and flexible automated tools for data analysis and science exploitation. We present the Galaxy Spectra Network/GaSNet-II, a supervised multi-network deep learning tool for spectra classification and redshift prediction. GaSNet-II can be trained to identify a customized number of classes and optimize the redshift predictions for classified objects in each of them. It also provides redshift errors, using a network-of-networks that reproduces a Monte Carlo test on each spectrum, by randomizing their weight initialization. As a demonstration of the capability of the deep learning pipeline, we use 260k Sloan Digital Sky Survey spectra from Data Release 16, separated into 13 classes including 140k galactic, and 120k extragalactic objects. GaSNet-II achieves 92.4% average classification accuracy over the 13 classes (larger than 90% for the majority of them), and an average redshift error of approximately 0.23% for galaxies and 2.1% for quasars. We further train/test the same pipeline to classify spectra and predict redshifts for a sample of 200k 4MOST mock spectra and 21k publicly released DESI spectra. On 4MOST mock data, we reach 93.4% accuracy in 10-class classification and an average redshift error of 0.55% for galaxies and 0.3% for active galactic nuclei. On DESI data, we reach 96% accuracy in (star/galaxy/quasar only) classification and an average redshift error of 2.8% for galaxies and 4.8% for quasars, despite the small sample size available. GaSNet-II can process ~40k spectra in less than one minute, on a normal Desktop GPU. This makes the pipeline particularly suitable for real-time analyses of Stage-IV survey observations and an ideal tool for feedback loops aimed at night-by-night survey strategy optimization.

astro-ph.IM

On the general nature of 21cm-Lyman-$\alpha$ emitters cross-correlations during reionisation

We explore how the characteristics of the cross-correlation functions between the 21cm emission from the spin-flip transition of neutral hydrogen (HI) and early Lyman-$\alpha$ (Ly$\alpha$) radiation emitting galaxies (Ly$\alpha$ emitters, LAEs) depend on the reionisation history and topology and the simulated volume. For this purpose, we develop an analytic expression for the 21cm-LAE cross-correlation function and compare it to results derived from different Astraeus and 21cmFAST reionisation simulations covering a physically plausible range of scenarios where either low-mass ($<10^{9.5}M_\odot$) or massive ($>10^{9.5}M_\odot$) galaxies drive reionisation. Our key findings are: (i) the negative small-scale ($<2$ cMpc) cross-correlation amplitude scales with the intergalactic medium's (IGM) average HI fraction ($\langle\chi_\mathrm{HI}\rangle$) and spin-temperature weighted overdensity in neutral regions ($\langle1+\delta\rangle_\mathrm{HI}$); (ii) the inversion point of the cross-correlation function traces the peak of the size distribution of ionised regions around LAEs; (iii) the cross-correlation amplitude at small scales is sensitive to the reionisation topology, with its anti-correlation or correlation decreasing the stronger the ionising emissivity of the underlying galaxy population is correlated to the cosmic web gas distribution (i.e. the more low-mass galaxies drive reionisation); (iv) the required simulation volume to not underpredict the 21cm-LAE anti-correlation amplitude when the cross-correlation is derived via the cross-power spectrum rises as the size of ionised regions and their variance increases. Our analytic expression can serve two purposes: to test whether simulation volumes are sufficiently large, and to act as a fitting function when cross-correlating future 21cm signal Square Kilometre Array and LAE galaxy observations.

astro-ph.CO