Searcharxiv⌕ Search

arXiv subjects

Matthew Ho

Publications and source records attributed to Matthew Ho.

At least 37 records · Page 2Linked to original sources

Learning the Universe: physically-motivated priors for dust attenuation curves

Understanding the impact of dust on the spectral energy distributions (SEDs) of galaxies is crucial for inferring their physical properties and for studying the nature of interstellar dust. We analyze dust attenuation curves for $\sim 6400$ galaxies ($M_{\star} \sim 10^9 - 10^{11.5}\,M_{\odot}$) at $z=0.07$ in the IllustrisTNG50 and TNG100 simulations. Using radiative transfer post-processing, we generate synthetic attenuation curves and fit them with a parametric model that captures known extinction and attenuation laws (e.g., Calzetti, MW, SMC, LMC) and more exotic forms. We present the distributions of the best-fitting parameters: UV slope ($c_1$), optical-to-NIR slope ($c_2$), FUV slope ($c_3$), 2175 Angstrom bump strength ($c_4$), and normalization ($A_{\rm V}$). Key correlations emerge between $A_{\rm V}$ and the star formation rate surface density $Σ_{\rm SFR}$, as well as the UV slope $c_1$. The UV and FUV slopes ($c_1, c_3$) and the bump strength and visual attenuation ($c_4, A_{\rm V}$) exhibit robust internal correlations. Using these insights from simulations, we provide a set of scaling relations that predict a galaxy's median (averaged over line of sight) dust attenuation curve based solely on its $Σ_{\rm SFR}$ and/or $A_{\rm V}$. These predictions agree well with observed attenuation curves from the GALEX-SDSS-WISE Legacy Catalog despite minor differences in bump strength. This study delivers the most comprehensive library of synthetic attenuation curves for local galaxies, providing a foundation for physically motivated priors in SED fitting and galaxy inference studies, such as those performed as part of the Learning the Universe Collaboration.

astro-ph.GA↗

Learning the Universe: $3\ h^{-1}{\rm Gpc}$ Tests of a Field Level $N$-body Simulation Emulator

We apply and test a field-level emulator for non-linear cosmic structure formation in a volume matching next-generation surveys. Inferring the cosmological parameters and initial conditions from which the particular galaxy distribution of our Universe was seeded can be achieved by comparing simulated data to observational data. Previous work has focused on building accelerated forward models that efficiently mimic these simulations. One of these accelerated forward models uses machine learning to apply a non-linear correction to the linear $z=0$ Zeldovich approximation (ZA) fields, closely matching the cosmological statistics in the $N$-body simulation. This emulator was trained and tested at $(h^{-1}{\rm Gpc})^3$ volumes, although cosmological inference requires significantly larger volumes. We test this emulator at $(3\ h^{-1}{\rm Gpc})^3$ by comparing emulator outputs to $N$-body simulations for eight unique cosmologies. We consider several summary statistics, applied to both the raw particle fields and the dark matter (DM) haloes. We find that the power spectrum, bispectrum and wavelet statistics of the raw particle fields agree with the $N$-body simulations within ${\sim} 5 \%$ at most scales. For the haloes, we find a similar agreement between the emulator and the $N$-body for power spectrum and bispectrum, though a comparison of the stacked profiles of haloes shows that the emulator has slight errors in the positions of particles in the highly non-linear interior of the halo. At these large $(3\ h^{-1}{\rm Gpc})^3$ volumes, the emulator can create $z=0$ particle fields in a thousandth of the time required for $N$-body simulations and will be a useful tool for large-scale cosmological inference. This is a Learning the Universe publication.

astro-ph.CO↗

Bye-bye, Local-in-matter-density Bias: The Statistics of the Halo Field Are Poorly Determined by the Local Mass Density

Bias models relating the dark matter field to the spatial distribution of halos are widely used in current cosmological analyses. Many models predict halos purely from the local Eulerian matter density, yet bias models in perturbation theory require other local properties. We assess the validity of assuming that only the local dark matter density can be used to predict the number density of halos in a model-independent way and in the non-perturbative regime. Utilising $N$-body simulations, we study the properties of the halo counts field after spatial voxels with near-equal dark matter density have been permuted. If local-in-matter-density biasing were valid, the statistical properties of the permuted and un-permuted fields would be indistinguishable since both represent equally fair draws of the stochastic biasing model. If the Lagrangian radius is greater than approximately half the voxel size and for halos less massive than $\sim10^{15}\,h^{-1}{\rm\,M_\odot}$, we find the permuted halo field has a scale-dependent bias with greater than 25% more power on scales relevant for current surveys. These bias models remove small-scale power by not modelling correlations between neighbouring voxels, which substantially boosts large-scale power to conserve the field's total variance. This conclusion is robust to the choice of initial conditions and cosmology. Assuming local-in-matter-density halo biasing cannot, therefore, reproduce the distribution of halos across a large range of scales and halo masses, no matter how complex the model. One must either allow the biasing to be a function of other quantities and/or remove the assumption that neighbouring voxels are statistically independent.

astro-ph.CO↗

Proof Flow: Preliminary Study on Generative Flow Network Language Model Tuning for Formal Reasoning

Reasoning is a fundamental substrate for solving novel and complex problems. Deliberate efforts in learning and developing frameworks around System 2 reasoning have made great strides, yet problems of sufficient complexity remain largely out of reach for open models. To address this gap, we examine the potential of Generative Flow Networks as a fine-tuning method for LLMs to unlock advanced reasoning capabilities. In this paper, we present a proof of concept in the domain of formal reasoning, specifically in the Neural Theorem Proving (NTP) setting, where proofs specified in a formal language such as Lean can be deterministically and objectively verified. Unlike classical reward-maximization reinforcement learning, which frequently over-exploits high-reward actions and fails to effectively explore the state space, GFlowNets have emerged as a promising approach for sampling compositional objects, improving generalization, and enabling models to maintain diverse hypotheses. Our early results demonstrate GFlowNet fine-tuning's potential for enhancing model performance in a search setting, which is especially relevant given the paradigm shift towards inference time compute scaling and "thinking slowly."

cs.CL↗

CHARM: Creating Halos with Auto-Regressive Multi-stage networks

To maximize the amount of information extracted from cosmological datasets, simulations that accurately represent these observations are necessary. However, traditional simulations that evolve particles under gravity by estimating particle-particle interactions (N-body simulations) are computationally expensive and prohibitive to scale to the large volumes and resolutions necessary for the upcoming datasets. Moreover, modeling the distribution of galaxies typically involves identifying virialized dark matter halos, which is also a time- and memory-consuming process for large N-body simulations, further exacerbating the computational cost. In this study, we introduce CHARM, a novel method for creating mock halo catalogs by matching the spatial, mass, and velocity statistics of halos directly from the large-scale distribution of the dark matter density field. We develop multi-stage neural spline flow-based networks to learn this mapping at redshift z=0.5 directly with computationally cheaper low-resolution particle mesh simulations instead of relying on the high-resolution N-body simulations. We show that the mock halo catalogs and painted galaxy catalogs have the same statistical properties as obtained from $N$-body simulations in both real space and redshift space. Finally, we use these mock catalogs for cosmological inference using redshift-space galaxy power spectrum, bispectrum, and wavelet-based statistics using simulation-based inference, performing the first inference with accelerated forward model simulations and finding unbiased cosmological constraints with well-calibrated posteriors. The code was developed as part of the Simons Collaboration on Learning the Universe and is publicly available at \url{https://github.com/shivampcosmo/CHARM}.

astro-ph.CO↗

Inpainting Galaxy Counts onto N-Body Simulations over Multiple Cosmologies and Astrophysics

Cosmological hydrodynamical simulations, while the current state-of-the art methodology for generating theoretical predictions for the large scale structures of the Universe, are among the most expensive simulation tools, requiring upwards of 100 millions CPU hours per simulation. N-body simulations, which exclusively model dark matter and its purely gravitational interactions, represent a less resource-intensive alternative, however, they do not model galaxies, and as such cannot directly be compared to observations. In this study, we use conditional score-based models to learn a mapping from N-body to hydrodynamical simulations, specifically from dark matter density fields to the observable distribution of galaxies. We demonstrate that our model is capable of generating galaxy fields statistically consistent with hydrodynamical simulations at a fraction of the computational cost, and demonstrate our emulator is significantly more precise than traditional emulators over the scales 0.36 $h\ \text{Mpc}^{-1}$ $\leq$ k $\leq$ 3.88 $h\ \text{Mpc}^{-1}$.

astro-ph.CO↗

LtU-ILI: An All-in-One Framework for Implicit Inference in Astrophysics and Cosmology

This paper presents the Learning the Universe Implicit Likelihood Inference (LtU-ILI) pipeline, a codebase for rapid, user-friendly, and cutting-edge machine learning (ML) inference in astrophysics and cosmology. The pipeline includes software for implementing various neural architectures, training schemata, priors, and density estimators in a manner easily adaptable to any research workflow. It includes comprehensive validation metrics to assess posterior estimate coverage, enhancing the reliability of inferred results. Additionally, the pipeline is easily parallelizable and is designed for efficient exploration of modeling hyperparameters. To demonstrate its capabilities, we present real applications across a range of astrophysics and cosmology problems, such as: estimating galaxy cluster masses from X-ray photometry; inferring cosmology from matter power spectra and halo point clouds; characterizing progenitors in gravitational wave signals; capturing physical dust parameters from galaxy colors and luminosities; and establishing properties of semi-analytic models of galaxy formation. We also include exhaustive benchmarking and comparisons of all implemented methods as well as discussions about the challenges and pitfalls of ML inference in astronomical sciences. All code and examples are made publicly available at https://github.com/maho3/ltu-ili.

astro-ph.IM↗

Reliable computation by large-alphabet formulas in the presence of noise

We present two new positive results for reliable computation using formulas over physical alphabets of size $q > 2$. First, we show that for logical alphabets of size $\ell = q$ the threshold for denoising using gates subject to $q$-ary symmetric noise with error probability $\varepsilon$ is strictly larger than that for Boolean computation, and is possible as long as signals remain distinguishable, i.e. $ε< (q - 1) / q$, in the limit of large fan-in $k \rightarrow \infty$. We also determine the point at which generalized majority gates with bounded fan-in fail, and show in particular that reliable computation is possible for $ε< (q - 1) / (q (q + 1))$ in the case of $q$ prime and fan-in $k = 3$. Secondly, we provide an example where $\ell < q$, showing that reliable Boolean computation can be performed using $2$-input ternary logic gates subject to symmetric ternary noise of strength $\varepsilon < 1/6$ by using the additional alphabet element for error signaling.

cs.IT↗

Metallicity Dependence of Pressure-Regulated Feedback-Modulated Star Formation in the TIGRESS-NCR Simulation Suite

We present a new simulation suite for the star-forming interstellar medium (ISM) in galactic disks using the TIGRESS-NCR framework. Distinctive aspects of our simulation suite are: (1) sophisticated and comprehensive numerical treatments of essential physical processes including magnetohydrodynamics, self-gravity, and galactic differential rotation, as well as photochemistry, cooling, and heating coupled with ray-tracing UV radiation transfer and resolved supernova feedback and (2) wide parameter coverage including metallicity over $Z'\equiv Z/Z_\odot\sim0.1-3$, gas surface density $Σ_{\rm gas}\sim5-150 M_{\odot}{\rm pc^{-2}}$, and stellar surface density $Σ_{\rm star}\sim 1-50 M_{\odot}{\rm pc^{-2}}$. The range of emergent star formation rate surface density is $Σ_{\rm SFR}\sim 10^{-4}-0.5 M_{\odot}{\rm kpc^{-2}yr^{-1}}$ and ISM total midplane pressure is $P_{\rm tot}/k_B=10^3-10^6{\rm cm^{-3}K}$, with $P_{\rm tot}$ equal to the ISM weight $W$. For given $Σ_{\rm gas}$ and $Σ_{\rm star}$, we find $Σ_{\rm SFR} \propto Z'^{0.3}$. We provide an interpretation based on the pressure-regulated feedback-modulated (PRFM) star formation theory. We characterize feedback modulation in terms of the yield $Υ$, defined as the ratio of each stress to $Σ_{\rm SFR}$. The thermal feedback yield varies sensitively with both weight and metallicity as $Υ_{\rm th}\propto W^{-0.46}Z'^{-0.53}$, while the combined turbulent and magnetic feedback yield shows weaker dependence $Υ_{\rm turb+mag}\propto W^{-0.22}Z'^{-0.18}$. The reduction in $Σ_{\rm SFR}$ at low metallicity is due mainly to enhanced thermal feedback yield, resulting from reduced attenuation of UV radiation. With the metallicity-dependent calibrations we provide, PRFM theory can be used for a new subgrid star formation prescription in cosmological simulations where the ISM is unresolved.

astro-ph.GA↗

Quantum-inspired identification of complex cellular automata

Elementary cellular automata (ECA) present iconic examples of complex systems. Though described only by one-dimensional strings of binary cells evolving according to nearest-neighbour update rules, certain ECA rules manifest complex dynamics capable of universal computation. Yet, the classification of precisely which rules exhibit complex behaviour remains a significant challenge. Here we approach this question using tools from quantum stochastic modelling, where quantum statistical memory -- the memory required to model a stochastic process using a class of quantum machines -- can be used to quantify the structure of a stochastic process. By viewing ECA rules as transformations of stochastic patterns, we ask: Does an ECA generate structure as quantified by the quantum statistical memory, and if so, how quickly? We illustrate how the growth of this measure over time correctly distinguishes simple ECA from complex counterparts. Moreover, it provides a more refined means for quantitatively identifying complex ECAs -- providing a spectrum on which we can rank the complexity of ECA by the rate in which they generate structure.

quant-ph↗

Sensitivity Analysis of Simulation-Based Inference for Galaxy Clustering

Simulation-based inference (SBI) is a promising approach to leverage high fidelity cosmological simulations and extract information from the non-Gaussian, non-linear scales that cannot be modeled analytically. However, scaling SBI to the next generation of cosmological surveys faces the computational challenge of requiring a large number of accurate simulations over a wide range of cosmologies, while simultaneously encompassing large cosmological volumes at high resolution. This challenge can potentially be mitigated by balancing the accuracy and computational cost for different components of the the forward model while ensuring robust inference. To guide our steps in this, we perform a sensitivity analysis of SBI for galaxy clustering on various components of the cosmological simulations: gravity model, halo-finder and the galaxy-halo distribution models (halo-occupation distribution, HOD). We infer the $σ_8$ and $Ω_m$ using galaxy power spectrum multipoles and the bispectrum monopole assuming a galaxy number density expected from the luminous red galaxies observed using the Dark Energy Spectroscopy Instrument (DESI). We find that SBI is insensitive to changing gravity model between $N$-body simulations and particle mesh (PM) simulations. However, changing the halo-finder from friends-of-friends (FoF) to Rockstar can lead to biased estimate of $σ_8$ based on the bispectrum. For galaxy models, training SBI on more complex HOD leads to consistent inference for less complex HOD models, but SBI trained on simpler HOD models fails when applied to analyze data from a more complex HOD model. Based on our results, we discuss the outlook on cosmological simulations with a focus on applying SBI approaches to future galaxy surveys.

astro-ph.CO↗

Benchmarks and Explanations for Deep Learning Estimates of X-ray Galaxy Cluster Masses

We evaluate the effectiveness of deep learning (DL) models for reconstructing the masses of galaxy clusters using X-ray photometry data from next-generation surveys. We establish these constraints using a catalogue of realistic mock eROSITA X-ray observations which use hydrodynamical simulations to model realistic cluster morphology, background emission, telescope response, and AGN sources. Using bolometric X-ray photon maps as input, DL models achieve a predictive mass scatter of $σ_{\ln M_\mathrm{500c}} = 17.8\%$, a factor of two improvements on scalar observables such as richness $N_\mathrm{gal}$, 1D velocity dispersion $σ_\mathrm{v,1D}$, and photon count $N_\mathrm{phot}$ as well as a $32\%$ improvement upon idealised, volume-integrated measurements of the bolometric X-ray luminosity $L_X$. We then show that extending this model to handle multichannel X-ray photon maps, separated in low, medium, and high energy bands, further reduces the mass scatter to $16.2\%$. We also tested a multimodal DL model incorporating both dynamical and X-ray cluster probes and achieved marginal gains at a mass scatter of $15.9\%$. Finally, we conduct a quantitative interpretability study of our DL models and find that they greatly down-weight the importance of pixels in the centres of clusters and at the location of AGN sources, validating previous claims of DL modelling improvements and suggesting practical and theoretical benefits for using DL in X-ray mass inference.

astro-ph.CO↗

Information-Ordered Bottlenecks for Adaptive Semantic Compression

We present the information-ordered bottleneck (IOB), a neural layer designed to adaptively compress data into latent variables ordered by likelihood maximization. Without retraining, IOB nodes can be truncated at any bottleneck width, capturing the most crucial information in the first latent variables. Unifying several previous approaches, we show that IOBs achieve near-optimal compression for a given encoding architecture and can assign ordering to latent signals in a manner that is semantically meaningful. IOBs demonstrate a remarkable ability to compress embeddings of image and text data, leveraging the performance of SOTA architectures such as CNNs, transformers, and diffusion models. Moreover, we introduce a novel theory for estimating global intrinsic dimensionality with IOBs and show that they recover SOTA dimensionality estimates for complex synthetic data. Furthermore, we showcase the utility of these models for exploratory analysis through applications on heterogeneous datasets, enabling computer-aided discovery of dataset complexity.

cs.LG↗

Posterior Sampling of the Initial Conditions of the Universe from Non-linear Large Scale Structures using Score-Based Generative Models

Reconstructing the initial conditions of the universe is a key problem in cosmology. Methods based on simulating the forward evolution of the universe have provided a way to infer initial conditions consistent with present-day observations. However, due to the high complexity of the inference problem, these methods either fail to sample a distribution of possible initial density fields or require significant approximations in the simulation model to be tractable, potentially leading to biased results. In this work, we propose the use of score-based generative models to sample realizations of the early universe given present-day observations. We infer the initial density field of full high-resolution dark matter N-body simulations from the present-day density field and verify the quality of produced samples compared to the ground truth based on summary statistics. The proposed method is capable of providing plausible realizations of the early universe density field from the initial conditions posterior distribution marginalized over cosmological parameters and can sample orders of magnitude faster than current state-of-the-art methods.

astro-ph.CO↗

WikiWhy: Answering and Explaining Cause-and-Effect Questions

As large language models (LLMs) grow larger and more sophisticated, assessing their "reasoning" capabilities in natural language grows more challenging. Recent question answering (QA) benchmarks that attempt to assess reasoning are often limited by a narrow scope of covered situations and subject matters. We introduce WikiWhy, a QA dataset built around a novel auxiliary task: explaining why an answer is true in natural language. WikiWhy contains over 9,000 "why" question-answer-rationale triples, grounded on Wikipedia facts across a diverse set of topics. Each rationale is a set of supporting statements connecting the question to the answer. WikiWhy serves as a benchmark for the reasoning capabilities of LLMs because it demands rigorous explicit rationales for each answer to demonstrate the acquisition of implicit commonsense knowledge, which is unlikely to be easily memorized. GPT-3 baselines achieve only 38.7% human-evaluated correctness in the end-to-end answer & explain condition, leaving significant room for future improvements.

cs.CL↗

A Machine Learning Approach to Enhancing eROSITA Observations

The eROSITA X-ray telescope, launched in 2019, is predicted to observe roughly 100,000 galaxy clusters. Follow-up observations of these clusters from Chandra, for example, will be needed to resolve outstanding questions about galaxy cluster physics. Deep Chandra cluster observations are expensive and follow-up of every eROSITA cluster is infeasible, therefore, objects chosen for follow-up must be chosen with care. To address this, we have developed an algorithm for predicting longer duration, background-free observations based on mock eROSITA observations. We make use of the hydrodynamic cosmological simulation Magneticum, have simulated eROSITA instrument conditions using SIXTE, and have applied a novel convolutional neural network to output a deep Chandra-like "super observation" of each cluster in our simulation sample. Any follow-up merit assessment tool should be designed with a specific use case in mind; our model produces observations that accurately and precisely reproduce the cluster morphology, which is a critical ingredient for determining cluster dynamical state and core type. Our model will advance our understanding of galaxy clusters by improving follow-up selection and demonstrates that image-to-image deep learning algorithms are a viable method for simulating realistic follow-up observations.

astro-ph.CO↗

The Dynamical Mass of the Coma Cluster from Deep Learning

In 1933, Fritz Zwicky's famous investigations of the mass of the Coma cluster led him to infer the existence of dark matter \cite{1933AcHPh...6..110Z}. His fundamental discoveries have proven to be foundational to modern cosmology; as we now know such dark matter makes up 85\% of the matter and 25\% of the mass-energy content in the universe. Galaxy clusters like Coma are massive, complex systems of dark matter in addition to hot ionized gas and thousands of galaxies, and serve as excellent probes of the dark matter distribution. However, empirical studies show that the total mass of such systems remains elusive and difficult to precisely constrain. Here, we present new estimates for the dynamical mass of the Coma cluster based on Bayesian deep learning methodologies developed in recent years. Using our novel data-driven approach, we predict Coma's $\mthc$ mass to be $10^{15.10 \pm 0.15}\ \hmsun$ within a radius of $1.78 \pm 0.03\ h^{-1}\mathrm{Mpc}$ of its center. We show that our predictions are rigorous across multiple training datasets and statistically consistent with historical estimates of Coma's mass. This measurement reinforces our understanding of the dynamical state of the Coma cluster and advances rigorous analyses and verification methods for empirical applications of machine learning in astronomy.

astro-ph.CO↗

Building Trustworthy Machine Learning Models for Astronomy

Astronomy is entering an era of data-driven discovery, due in part to modern machine learning (ML) techniques enabling powerful new ways to interpret observations. This shift in our scientific approach requires us to consider whether we can trust the black box. Here, we overview methods for an often-overlooked step in the development of ML models: building community trust in the algorithms. Trust is an essential ingredient not just for creating more robust data analysis techniques, but also for building confidence within the astronomy community to embrace machine learning methods and results.

astro-ph.IM↗