SearcharxivSearch

arXiv subjects

Zarija Lukic

Publications and source records attributed to Zarija Lukic.

At least 19 recordsLinked to original sources

A Measurement of the Thermal and Ionization State of the IGM at $z < 0.5$

We apply a machine-learning-based inference method that exploits the joint Doppler parameter-column density (b-NHI) distribution from Lya forest decomposition to measure the thermal and ionization state of the intergalactic medium (IGM) in four redshift bins spanning z = 0.06 to 0.48, using 82 archival quasar spectra from the Cosmic Origin Spectrograph (COS) on board Hubble Space Telescope (HST). Our results show that the low-z IGM (z < 0.5) is extremely hot and nearly isothermal, with log(T0/K) = 4.45 (+0.08 / -0.12) [T0 = 28183 (+5700 / -6804) K] and gamma = 1.06 (+0.13 / -0.09) at z = 0.1. This temperature lies approx 7sigma (and 7 times) above the canonical prediction (log T0 approx 3.60, i.e. T0 ~ 4000 K, with gamma ~ 1.6 at z = 0), where the IGM is expected to have cooled long after He II reionization. We also measure the hydrogen photoionization rate to be log (GammaHI/s^-1) = -13.70 (+0.10 / -0.08) at z = 0.1, which is about approx 4sigma below the range predicted by current UV-background synthesis models (approx -13.3). To investigate the discrepancy between these high temperatures and theoretical models, we assess the impact of small-scale turbulence. By exploring a parameter grid in turbulent velocity (vtur) and GammaHI, we find that a standard IGM thermal and ionization state combined with unresolved turbulence of vtur simeq 15 km s^-1 can successfully reproduce the observed line widths at z = 0.1. Comparisons with high-resolution Space Telescope Imaging Spectrograph (STIS) expanded data indicate that the observed line widths are unlikely to be caused by instrumental resolution effects. Our findings suggest that either new heating mechanisms or unresolved turbulence are required to explain the unexpectedly broad Lya lines observed in the low-z IGM.

astro-ph.CO

FFCz: Fast Fourier Correction for Spectrum-Preserving Lossy Compression of Scientific Data

This paper introduces a novel technique to preserve spectral features in lossy compression based on a novel fast Fourier correction algorithm\added{ for regular-grid data}. Preserving both spatial and frequency representations of data is crucial for applications such as cosmology, turbulent combustion, and X-ray diffraction, where spatial and frequency views provide complementary scientific insights. In particular, many analysis tasks rely on frequency-domain representations to capture key features, including the power spectrum of cosmology simulations, the turbulent energy spectrum in combustion, and diffraction patterns in reciprocal space for ptychography. However, existing compression methods guarantee accuracy only in the spatial domain while disregarding the frequency domain. To address this limitation, we propose an algorithm that corrects the errors produced by off-the-shelf ``base'' compressors such as SZ3, ZFP, and SPERR, thereby preserving both spatial and frequency representations by bounding errors in both domains. By expressing frequency-domain errors as linear combinations of spatial-domain errors, we derive a region that jointly bounds errors in both domains. Given as input the spatial errors from a base compressor and user-defined error bounds in the spatial and frequency domains, we iteratively project the spatial error vector onto the regions defined by the spatial and frequency constraints until it lies within their intersection. We further accelerate the algorithm using GPU parallelism to achieve practical performance. We validate our approach with datasets from cosmology simulations, X-ray diffraction, combustion simulation, and electroencephalography demonstrating its effectiveness in preserving critical scientific information in both spatial and frequency domains.

cs.DC

SuperBench: A Super-Resolution Benchmark Dataset for Scientific Machine Learning

Super-resolution (SR) techniques aim to enhance data resolution, enabling the retrieval of finer details, and improving the overall quality and fidelity of the data representation. There is growing interest in applying SR methods to complex spatiotemporal systems within the Scientific Machine Learning (SciML) community, with the hope of accelerating numerical simulations and/or improving forecasts in weather, climate, and related areas. However, the lack of standardized benchmark datasets for comparing and validating SR methods hinders progress and adoption in SciML. To address this, we introduce SuperBench, the first benchmark dataset featuring high-resolution datasets, including data from fluid flows, cosmology, and weather. Here, we focus on validating spatial SR performance from data-centric and physics-preserved perspectives, as well as assessing robustness to data degradation tasks. While deep learning-based SR methods (developed in the computer vision community) excel on certain tasks, despite relatively limited prior physics information, we identify limitations of these methods in accurately capturing intricate fine-scale features and preserving fundamental physical properties and constraints in scientific data. These shortcomings highlight the importance and subtlety of incorporating domain knowledge into ML models. We anticipate that SuperBench will help to advance SR methods for science.

cs.CV

Differentiable Cosmological Hydrodynamics for Field-Level Inference and High Dimensional Parameter Constraints

Hydrodynamical simulations are the most accurate way to model structure formation in the universe, but they often involve a large number of astrophysical parameters modeling subgrid physics, in addition to cosmological parameters. This results in a high-dimensional space that is difficult to jointly constrain using traditional statistical methods due to prohibitive computational costs. To address this, we present a fully differentiable approach for cosmological hydrodynamical simulations and a proof-of-concept implementation, diffhydro. By back-propagating through an upwind finite volume scheme for solving the Euler Equations jointly with a dark matter particle-mesh method for Poisson equation, we are able to efficiently evaluate derivatives of the output baryonic fields with respect to input density and model parameters. Importantly, we demonstrate how to differentiate through stochastically sampled discrete random variables, which frequently appear in subgrid models. We use this framework to rapidly sample sub-grid physics and cosmological parameters as well as perform field level inference of initial conditions using high dimensional optimization techniques. Our code is implemented in JAX (python), allowing easy code development and GPU acceleration.

astro-ph.CO

The ACCEL2 Project: Precision Measurements of EFT Parameters and BAO Peak Shifts for the Lyman-$α$ Forest

We present precision measurements of the bias parameters of the one-loop power spectrum model of the Lyman-alpha (Lya) forest, derived within the effective field theory of large-scale structure (EFT). We fit our model to the three-dimensional flux power spectrum measured from the ACCEL2 hydrodynamic simulations. The EFT model fits the data with an accuracy of below 2 percent up to a wavenumber of k = 2 h/Mpc. Further, we analytically derive how non-linearities in the three-dimensional clustering of the Lya forest introduce biases in measurements of the Baryon Acoustic Oscillations (BAO) scaling parameters in radial and transverse directions. From our EFT parameter measurements, we obtain a theoretical error budget of -0.2 (-0.3) percent for the radial (transverse) parameters at redshift two. This corresponds to a shift of -0.3 (0.1) percent for the isotropic (anisotropic) distance measurements. We provide an estimate for the shift of the BAO peak for Lya-quasar cross-correlation measurements assuming analytical and simulation-based scaling relations for the non-linear quasar bias parameters resulting in a shift of -0.2 (-0.1) percent for the radial (transverse) dilation parameters, respectively. This analysis emphasizes the robustness of Lya forest BAO measurements to the theory modeling. We provide informative priors and an error budget for measuring the BAO feature -- a key science driver of the currently observing Dark Energy Spectroscopic Instrument (DESI). Our work paves the way for full-shape cosmological analyses of Lya forest data from DESI and upcoming surveys such as the Prime Focus Spectrograph, WEAVE-QSO, and 4MOST.

astro-ph.CO

Maximum A Posteriori Ly-alpha Estimator (MAPLE): Band-power and covariance estimation of the 3D Ly-alpha forest power spectrum

We present a novel maximum a posteriori estimator to jointly estimate band-powers and the covariance of the three-dimensional power spectrum (P3D) of Lyman-alpha forest flux fluctuations, called MAPLE. Our Wiener-filter based algorithm reconstructs a window-deconvolved P3D in the presence of complex survey geometries typical for Lyman-alpha surveys that are sparsely sampled transverse to and densely sampled along the line-of-sight. We demonstrate our method on idealized Gaussian random fields with two selection functions: (i) a sparse sampling of 30 background sources per square degree designed to emulate the currently observing the Dark Energy Spectroscopic Instrument (DESI); (ii) a dense sampling of 900 background sources per square degree emulating the upcoming Prime Focus Spectrograph Galaxy Evolution Survey. Our proof-of-principle shows promise, especially since the algorithm can be extended to marginalize jointly over nuisance parameters and contaminants, i.e.offsets introduced by continuum fitting. Our code is implemented in JAX and is publicly available on GitHub.

astro-ph.CO

Measurements of the Thermal and Ionization State of the Intergalactic Medium during the Cosmic Afternoon

We perform the first measurement of the thermal and ionization state of the intergalactic medium (IGM) across 0.9 < z < 1.5 using 301 \lya absorption lines fitted from 12 HST STIS quasar spectra, with a total pathlength of Δz=2.1. We employ the machine-learning-based inference method that uses joint b-N distributions obtained from \lyaf decomposition. Our results show that the HI photoionization rates, Γ, are in good agreement with the recent UV background synthesis models, with \log (Γ/s^{-1})={-11.79}^{0.18}_{-0.15}, -11.98}^{0.09}_{-0.09}, and {-12.32}^{0.10}_{-0.12} at z=1.4, 1.2, and 1 respectively. We obtain the IGM temperature at the mean density, T_0, and the adiabatic index, γ, as [\log (T_0/K), γ]= [{4.13}^{+0.12}_{-0.10}, {1.34}^{+0.10}_{-0.15}], [{3.79}^{+0.11}_{-0.11}, {1.70}^{+0.09}_{-0.09}] and [{4.12}^{+0.15}_{-0.25}, {1.34}^{+0.21}_{-0.26}] at z=1.4, 1.2 and 1 respectively. Our measurements of T_0 at z=1.4 and 1.2 are consistent with the expected trend from z<3 temperature measurements as well as theoretical expectations that, in the absence of any non-standard heating, the IGM should cool down after HeII reionization. Whereas, our T_0 measurements at z=1 show unexpectedly high IGM temperature. However, because of the relatively large uncertainty in these measurements of the order of ΔT_0~5000 K, mostly emanating from the limited redshift path length of available data in these bins, we can not definitively conclude whether the IGM cools down at z<1.5. Lastly, we generate a mock dataset to test the constraining power of future measurement with larger datasets. The results demonstrate that, with redshift pathlength Δz \sim 2 for each redshift bin, three times the current dataset, we can constrain the T_0 of IGM within 1500K. Such precision would be sufficient to conclusively constrain the history of IGM thermal evolution at z < 1.5.

astro-ph.CO

The Impact of the WHIM on the IGM Thermal State Determined from the Low-$z$ Lyman-$α$ Forest

At $z \lesssim 1$, shock heating caused by large-scale velocity flows and possibly violent feedback from galaxy formation, converts a significant fraction of the cool gas ($T\sim 10^4$ K) in the intergalactic medium (IGM) into warm-hot phase (WHIM) with $T >10^5$K, resulting in a significant deviation from the previously tight power-law IGM temperature-density relationship, $T=T_0 (ρ/ {\barρ})^{γ-1}$. This study explores the impact of the WHIM on measurements of the low-$z$ IGM thermal state, $[T_0,γ]$, based on the $b$-$N_{H I}$ distribution of the Lyman-$α$ forest. Exploiting a machine learning-enabled simulation-based inference method trained on Nyx hydrodynamical simulations, we demonstrate that [$T_0$, $γ$] can still be reliably measured from the $b$-$N_{H I}$ distribution at $z=0.1$, notwithstanding the substantial WHIM in the IGM. To investigate the effects of different feedback, we apply this inference methodology to mock spectra derived from the IllustrisTNG and Illustris simulations at $z=0.1$. The results suggest that the underlying $[T_0,γ]$ of both simulations can be recovered with biases as low as $|Δ\log(T_0/\text{K})| \lesssim 0.05$ dex, $|Δγ| \lesssim 0.1$, smaller than the precision of a typical measurement. Given the large differences in the volume-weighted WHIM fractions between the three simulations (Illustris 38\%, IllustrisTNG 10\%, Nyx 4\%) we conclude that the $b$-$N_{H I}$ distribution is not sensitive to the WHIM under realistic conditions. Finally, we investigate the physical properties of the detectable Lyman-$α$ absorbers, and discover that although their $T$ and $Δ$ distributions remain mostly unaffected by feedback, they are correlated with the photoionization rate used in the simulation.

astro-ph.CO

Measuring the thermal and ionization state of the low-$z$ IGM using likelihood free inference

We present a new approach to measure the power-law temperature density relationship $T=T_0 (ρ/ \barρ)^{γ-1}$ and the UV background photoionization rate $Γ_{\rm HI}$ of the IGM based on the Voigt profile decomposition of the Ly$α$ forest into a set of discrete absorption lines with Doppler parameter $b$ and the neutral hydrogen column density $N_{\rm HI}$. Previous work demonstrated that the shape of the $b$-$N_{\rm HI}$ distribution is sensitive to the IGM thermal parameters $T_0$ and $γ$, whereas our new inference algorithm also takes into account the normalization of the distribution, i.e. the line-density d$N$/d$z$, and we demonstrate that precise constraints can also be obtained on $Γ_{\rm HI}$. We use density-estimation likelihood-free inference (DELFI) to emulate the dependence of the $b$-$N_{\rm HI}$ distribution on IGM parameters trained on an ensemble of 624 Nyx hydrodynamical simulations at $z = 0.1$, which we combine with a Gaussian process emulator of the normalization. To demonstrate the efficacy of this approach, we generate hundreds of realizations of realistic mock HST/COS datasets, each comprising 34 quasar sightlines, and forward model the noise and resolution to match the real data. We use this large ensemble of mocks to extensively test our inference and empirically demonstrate that our posterior distributions are robust. Our analysis shows that by applying our new approach to existing Ly$α$ forest spectra at $z\simeq 0.1$, one can measure the thermal and ionization state of the IGM with very high precision ($σ_{\log T_0} \sim 0.08$ dex, $σ_γ\sim 0.06$, and $σ_{\log Γ_{\rm HI}} \sim 0.07$ dex).

astro-ph.CO

Snowmass2021 Cosmic Frontier White Paper: Prospects for obtaining Dark Matter Constraints with DESI

Despite efforts over several decades, direct-detection experiments have not yet led to the discovery of the dark matter (DM) particle. This has led to increasing interest in alternatives to the Lambda CDM (LCDM) paradigm and alternative DM scenarios (including fuzzy DM, warm DM, self-interacting DM, etc.). In many of these scenarios, DM particles cannot be detected directly and constraints on their properties can ONLY be arrived at using astrophysical observations. The Dark Energy Spectroscopic Instrument (DESI) is currently one of the most powerful instruments for wide-field surveys. The synergy of DESI with ESA's Gaia satellite and future observing facilities will yield datasets of unprecedented size and coverage that will enable constraints on DM over a wide range of physical and mass scales and across redshifts. DESI will obtain spectra of the Lyman-alpha forest out to z~5 by detecting about 1 million QSO spectra that will put constraints on clustering of the low-density intergalactic gas and DM halos at high redshift. DESI will obtain radial velocities of 10 million stars in the Milky Way (MW) and Local Group satellites enabling us to constrain their global DM distributions, as well as the DM distribution on smaller scales. The paradigm of cosmological structure formation has been extensively tested with simulations. However, the majority of simulations to date have focused on collisionless CDM. Simulations with alternatives to CDM have recently been gaining ground but are still in their infancy. While there are numerous publicly available large-box and zoom-in simulations in the LCDM framework, there are no comparable publicly available WDM, SIDM, FDM simulations. DOE support for a public simulation suite will enable a more cohesive community effort to compare observations from DESI (and other surveys) with numerical predictions and will greatly impact DM science.

astro-ph.CO

Accelerating Parallel Write via Deeply Integrating Predictive Lossy Compression with HDF5

Lossy compression is one of the most efficient solutions to reduce storage overhead and improve I/O performance for HPC applications. However, existing parallel I/O libraries cannot fully utilize lossy compression to accelerate parallel write due to the lack of deep understanding on compression-write performance. To this end, we propose to deeply integrate predictive lossy compression with HDF5 to significantly improve the parallel-write performance. Specifically, we propose analytical models to predict the time of compression and parallel write before the actual compression to enable compression-write overlapping. We also introduce an extra space in the process to handle possible data overflows resulting from prediction uncertainty in compression ratios. Moreover, we propose an optimization to reorder the compression tasks to increase the overlapping efficiency. Experiments with up to 4,096 cores from Summit show that our solution improves the write performance by up to 4.5X and 2.9X over the non-compression and lossy compression solutions, respectively, with only 1.5% storage overhead (compared to original data) on two real-world HPC applications.

cs.DC

Mining for Strong Gravitational Lenses with Self-supervised Learning

We employ self-supervised representation learning to distill information from 76 million galaxy images from the Dark Energy Spectroscopic Instrument Legacy Imaging Surveys' Data Release 9. Targeting the identification of new strong gravitational lens candidates, we first create a rapid similarity search tool to discover new strong lenses given only a single labelled example. We then show how training a simple linear classifier on the self-supervised representations, requiring only a few minutes on a CPU, can automatically classify strong lenses with great efficiency. We present 1192 new strong lens candidates that we identified through a brief visual identification campaign, and release an interactive web-based similarity search tool and the top network predictions to facilitate crowd-sourcing rapid discovery of additional strong gravitational lenses and other rare objects: https://github.com/georgestein/ssl-legacysurvey.

astro-ph.IM

Self-supervised similarity search for large scientific datasets

We present the use of self-supervised learning to explore and exploit large unlabeled datasets. Focusing on 42 million galaxy images from the latest data release of the Dark Energy Spectroscopic Instrument (DESI) Legacy Imaging Surveys, we first train a self-supervised model to distill low-dimensional representations that are robust to symmetries, uncertainties, and noise in each image. We then use the representations to construct and publicly release an interactive semantic similarity search tool. We demonstrate how our tool can be used to rapidly discover rare objects given only a single example, increase the speed of crowd-sourcing campaigns, and construct and improve training sets for supervised applications. While we focus on images from sky surveys, the technique is straightforward to apply to any scientific dataset of any dimensionality. The similarity search web app can be found at https://github.com/georgestein/galaxy_search

astro-ph.IM

Report from the Tri-Agency Cosmological Simulation Task Force

The Tri-Agency Cosmological Simulations (TACS) Task Force was formed when Program Managers from the Department of Energy (DOE), the National Aeronautics and Space Administration (NASA), and the National Science Foundation (NSF) expressed an interest in receiving input into the cosmological simulations landscape related to the upcoming DOE/NSF Vera Rubin Observatory (Rubin), NASA/ESA's Euclid, and NASA's Wide Field Infrared Survey Telescope (WFIRST). The Co-Chairs of TACS, Katrin Heitmann and Alina Kiessling, invited community scientists from the USA and Europe who are each subject matter experts and are also members of one or more of the surveys to contribute. The following report represents the input from TACS that was delivered to the Agencies in December 2018.

astro-ph.CO

Cosmic Inference: Constraining Parameters With Observations and Highly Limited Number of Simulations

Cosmological probes pose an inverse problem where the measurement result is obtained through observations, and the objective is to infer values of model parameters which characterize the underlying physical system -- our Universe. Modern cosmological probes increasingly rely on measurements of the small-scale structure, and the only way to accurately model physical behavior on those scales, roughly 65 Mpc/h or smaller, is via expensive numerical simulations. In this paper, we provide a detailed description of a novel statistical framework for obtaining accurate parameter constraints by combining observations with a very limited number of cosmological simulations. The proposed framework utilizes multi-output Gaussian process emulators that are adaptively constructed using Bayesian optimization methods. We compare several approaches for constructing multi-output emulators that enable us to take possible inter-output correlations into account while maintaining the efficiency needed for inference. Using Lyman alpha forest flux power spectrum, we demonstrate that our adaptive approach requires considerably fewer --- by a factor of a few in Lyman alpha P(k) case considered here --- simulations compared to the emulation based on Latin hypercube sampling, and that the method is more robust in reconstructing parameters and their Bayesian credible intervals.

astro-ph.IM

Mapping quasar light echoes in 3D with Lyα forest tomography

The intense radiation emitted by luminous quasars dramatically alters the ionization state of their surrounding IGM. This so-called proximity effect extends out to tens of Mpc, and manifests as large coherent regions of enhanced Lyman-$α$ (Ly$α$) forest transmission in absorption spectra of background sightlines. Here we present a novel method based on Ly$α$ forest tomography, which is capable of mapping these quasar `light echoes' in three dimensions. Using a dense grid (10-100) of faint ($m_r\approx24.7\,\mathrm{mag}$) background galaxies as absorption probes, one can measure the ionization state of the IGM in the vicinity of a foreground quasar, yielding detailed information about the quasar's radiative history and emission geometry. An end-to-end analysis - combining cosmological hydrodynamical simulations post-processed with a quasar emission model, realistic estimates of galaxy number densities, and instrument + telescope throughput - is conducted to explore the feasibility of detecting quasar light echoes. We present a new fully Bayesian statistical method that allows one to reconstruct quasar light echoes from thousands of individual low S/N transmission measurements. Armed with this machinery, we undertake an exhaustive parameter study and show that light echoes can be convincingly detected for luminous ($M_{1450} < -27.5\,\mathrm{mag}$ corresponding to $m_{1450} < 18.4\,\mathrm{mag}$ at $z\simeq 3.6$) quasars at redshifts $3 5$ is sufficient, requiring three hour integrations using existing instruments on 8m class telescopes.

astro-ph.GA

Measuring alignments between galaxies and the cosmic web at $z \sim 2-3$ using IGM tomography

Many galaxy formation models predict alignments between galaxy spin and the cosmic web (i.e. the directions of filaments and sheets), leading to intrinsic alignment between galaxies that creates a systematic error in weak lensing measurements. These effects are often predicted to be stronger at high-redshifts ($z\gtrsim1$) that are inaccessible to massive galaxy surveys on foreseeable instrumentation, but IGM tomography of the Ly$α$ forest from closely-spaced quasars and galaxies is starting to measure the $z\sim2-3$ cosmic web with the requisite fidelity. Using mock surveys from hydrodynamical simulations, we examine the utility of this technique, in conjunction with coeval galaxy samples, to measure alignment between galaxies and the cosmic web at $z\sim2.5$. We show that IGM tomography surveys with $\lesssim5$ $h^{-1}$ Mpc sightline spacing can accurately recover the eigenvectors of the tidal tensor, which we use to define the directions of the cosmic web. For galaxy spins and shapes, we use a model parametrized by the alignment strength, $Δ\langle\cosθ\rangle$, with respect to the tidal tensor eigenvectors from the underlying density field, and also consider observational effects such as errors in the galaxy position angle, inclination, and redshift. Measurements using the upcoming $\sim1\,\mathrm{deg}^2$ CLAMATO tomographic survey and 600 coeval zCOSMOS-Deep galaxies should place $3σ$ limits on extreme alignment models with $Δ\langle\cosθ\rangle\sim0.1$, but much larger surveys encompassing $>10,000$ galaxies, such as Subaru PFS, will be required to constrain models with $Δ\langle\cosθ\rangle\sim0.03$. These measurements will constrain models of galaxy-cosmic web alignment and test tidal torque theory at $z\sim2$, improving our understanding of the redshift dependence of galaxy-cosmic web alignment and the physics of intrinsic alignments.

astro-ph.CO

HACC: Simulating Sky Surveys on State-of-the-Art Supercomputing Architectures

Current and future surveys of large-scale cosmic structure are associated with a massive and complex datastream to study, characterize, and ultimately understand the physics behind the two major components of the 'Dark Universe', dark energy and dark matter. In addition, the surveys also probe primordial perturbations and carry out fundamental measurements, such as determining the sum of neutrino masses. Large-scale simulations of structure formation in the Universe play a critical role in the interpretation of the data and extraction of the physics of interest. Just as survey instruments continue to grow in size and complexity, so do the supercomputers that enable these simulations. Here we report on HACC (Hardware/Hybrid Accelerated Cosmology Code), a recently developed and evolving cosmology N-body code framework, designed to run efficiently on diverse computing architectures and to scale to millions of cores and beyond. HACC can run on all current supercomputer architectures and supports a variety of programming models and algorithms. It has been demonstrated at scale on Cell- and GPU-accelerated systems, standard multi-core node clusters, and Blue Gene systems. HACC's design allows for ease of portability, and at the same time, high levels of sustained performance on the fastest supercomputers available. We present a description of the design philosophy of HACC, the underlying algorithms and code structure, and outline implementation details for several specific architectures. We show selected accuracy and performance results from some of the largest high resolution cosmological simulations so far performed, including benchmarks evolving more than 3.6 trillion particles.

astro-ph.IM