SearcharxivSearch

arXiv subjects

Benjamin M. Boyd

Publications and source records attributed to Benjamin M. Boyd.

9 recordsLinked to original sources

Exploiting weight-space symmetries for approximating curvature

Many machine learning techniques rely on approximating a loss function's curvature, but this is notoriously hard to do at the scale of modern deep networks. Surprisingly, no previous work has exploited the curvature constraints that arise from well known weight-space symmetries in loss landscapes. By analytically averaging over group actions that leave the loss invariant, we construct structured Hessian approximations from single gradients that can be tractably estimated, stored, and inverted. The choice of user-specified symmetry group directly governs the trade-off between approximation accuracy and computational cost. Moreover, our framework provides a unifying theoretical lens for viewing existing methods; in particular, a specific choice of symmetry group recovers Shampoo/Muon-like curvature estimates. We validate our method on a range of network architectures, and deploy it to second-order optimization benchmarks, including a small language model. Our curvature estimation framework might find applications in other machine learning problems such as uncertainty estimation, continual learning, compression/pruning, training data attribution, and more.

cs.LG

On the origin of the environmental step: A BayeSN view of the ZTF SN Ia DR2

Astrophysical variabilities of Type Ia supernovae (SNe Ia), such as their link with their birth environment, are now one of the leading sources of systematic uncertainties on the measurement of the dark energy equation-of-state parameter $w$. Population studies of SNe Ia, using large samples, give precious insights into these variabilities. We analyse a volume-limited subsample of the ZTF SN Ia DR2 with BayeSN, a hierarchical Bayesian model for SN Ia SEDs. We investigate the distributions of SN Ia light curve parameters and their link with SN environment. Using a new training of BayeSN released in a companion paper, we find a smaller scatter of Hubble residuals compared to SALT. We then investigate the magnitude step, which accounts for the correlation between SN Ia standardised absolute magnitude and host environments. We find a posteriori steps of $0.103\pm0.010$ mag (a $10.1\sigma$ difference from 0) when using global stellar mass as an environmental proxy, and $0.086\pm0.010$ mag ($8.3\sigma$) when using local colour, in accordance with steps computed using SALT light curve fits. This confirms that the large step seen in the ZTF SN Ia DR2 data was not due to the SALT fit or the associated standardisation process. We then investigate the origin of the step, using a BayeSN model which accounts for both an intrinsic magnitude step and differing dust properties with the SN environment. We find a $0.103\pm0.018$ mag ($5.6\sigma$) step in global mass and a $0.085\pm0.019$ mag ($4.5\sigma$) step in local colour. The means of the $R_V$ distribution are similar between different host environments, with $\Delta\mathbb{E}(R_V)\leq0.2$ across all environment proxies, with significances ranging from $0.6\sigma$ to $1.2\sigma$. This is a strong signal of the existence of an intrinsic dependence of SN Ia absolute magnitude on environment.

astro-ph.CO

FlowSN: Neural Simulation-Based Inference under Realistic Selection Effects applied to Supernova Cosmology

We present FlowSN, a statistical framework using simulation-based inference (SBI) with normalising flows to account for selection effects in observational astronomy. Failure to account for selection effects can lead to biased inference on global parameters. An example is Malmquist bias, where detection limits result in a sample skewed towards brighter objects. In Type Ia supernova (SN Ia) cosmology, these selection effects can systematically shift the inferred posterior distributions of cosmological parameters, necessitating the development of robust statistical frameworks to account for the biases. SBI enables us to implicitly learn probability distributions that are analytically intractable to calculate. In this work, we introduce a novel approach that employs a normalising flow to learn the non-analytic selected SN likelihood for a given survey from forward simulations, independent of the assumed cosmological model. The resulting likelihood approximation is incorporated into a hierarchical Bayesian framework and posterior sampling is performed using Hamiltonian Monte Carlo to obtain constraints on cosmological parameters conditioned on the observed data. The modular learnt likelihood approximation can be reused without retraining to evaluate different cosmological models, providing a key advantage over other SBI approaches. We demonstrate the performance of this methodology by training and testing the SBI technique using realistic LSST-like SNANA simulations for the first time. Our FlowSN approach yields accurate posterior estimates on cosmological parameters, including the dark energy equation of state $w_0$, that are an order of magnitude less biased than those obtained with conventional techniques and also exhibit improved frequentist calibration.

astro-ph.CO

Attaining Spectral Energy Distributions With Sub-Percent Uncertainties: All-Sky DA White Dwarf Spectrophotometric Standard Stars For Large Telescopes And Surveys

We present a synopsis of the project to establish thirty-two new faint ($ 16.5 \leq V \leq 19.8 $) DA white dwarfs as spectrophotometric standards distributed over the whole sky. Our results validate the use of fully radiative pure hydrogen model fluxes for hot DA white dwarfs to predict the observed broadband fluxes from near ultraviolet through the near infrared to accuracies of a few parts per thousand. After fitting the line of sight reddenings simultaneously with the model spectral energy distributions of these stars against spectroscopic and multi-band photometric observations, we have shown that residuals have an rms of typically 0.4 percent. This indicates that the complications from interstellar dust extinction have been adequately mitigated. Our stars supplement the three brighter DA white dwarfs that define the flux scale of CALSPEC. The consequent photometric accuracy, their all sky coverage, and their brightness range that matches the dynamic range of large telescopes, constitutes an unprecedented ensemble of standard stars for both ground as well as space based use. This paper targets readers who may wish to use these as standard stars, and provides for them the essential content to understand their strengths and limitations, without traversing the technical details of analysis that are already captured in a series of papers since 2016. The narrative here describes the motivation, justification, and evolution of the analysis methods; the input data that constrain the modeling; as well as the stability of our results in the face of future improvements in models.

astro-ph.IM

DAmodel: Hierarchical Bayesian Modelling of DA White Dwarfs for Spectrophotometric Calibration

We use hierarchical Bayesian modelling to calibrate a network of 32 all-sky faint DA white dwarf (DA WD) spectrophotometric standards ($16.5 < V < 19.5$) alongside three CALSPEC standards, from 912 \r{A} to 32 $\mu$m. The framework is the first of its kind to jointly infer photometric zeropoints and WD parameters (surface gravity $\log g$, effective temperature $T_{\text{eff}}$, extinction $A_V$, dust relation parameter $R_V$) by simultaneously modelling both photometric and spectroscopic data. We model panchromatic Hubble Space Telescope Wide Field Camera 3 (HST/WFC3) UVIS and IR photometry, HST/STIS UV spectroscopy and ground-based optical spectroscopy to sub-percent precision. Photometric residuals for the sample are the lowest yet yielding $<0.004$ mag RMS on average from the UV to the NIR, achieved by jointly inferring time-dependent changes in system sensitivity and WFC3/IR count-rate nonlinearity. Our GPU-accelerated implementation enables efficient sampling via Hamiltonian Monte Carlo, critical for exploring the high-dimensional posterior space. The hierarchical nature of the model enables population analysis of intrinsic WD and dust parameters. Inferred spectral energy distributions from this model will be essential for calibrating the James Webb Space Telescope as well as next-generation surveys, including Vera Rubin Observatory's Legacy Survey of Space and Time and the Nancy Grace Roman Space Telescope.

astro-ph.IM

The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TB of Astronomical Scientific Data

We present the MULTIMODAL UNIVERSE, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, the MULTIMODAL UNIVERSE contains hundreds of millions of astronomical observations, constituting 100\,TB of multi-channel and hyper-spectral images, spectra, multivariate time series, as well as a wide variety of associated scientific measurements and "metadata". In addition, we include a range of benchmark tasks representative of standard practices for machine learning methods in astrophysics. This massive dataset will enable the development of large multi-modal models specifically targeted towards scientific applications. All codes used to compile the MULTIMODAL UNIVERSE and a description of how to access the data is available at https://github.com/MultimodalUniverse/MultimodalUniverse

astro-ph.IM

Accounting for Selection Effects in Supernova Cosmology with Simulation-Based Inference and Hierarchical Bayesian Modelling

Type Ia supernovae (SNe Ia) are thermonuclear exploding stars that can be used to put constraints on the nature of our universe. One challenge with population analyses of SNe Ia is Malmquist bias, where we preferentially observe the brighter SNe due to limitations of our telescopes. If untreated, this bias can propagate through to our posteriors on cosmological parameters. In this paper, we develop a novel technique of using a normalising flow to learn the non-analytical likelihood of observing a SN Ia for a given survey from simulations, that is independent of any cosmological model. The learnt likelihood is then used in a hierarchical Bayesian model with Hamiltonian Monte Carlo sampling to put constraints on different sets of cosmological parameters conditioned on the observed data. We verify this technique on toy model simulations finding excellent agreement with analytically-derived posteriors to within $1 \sigma$.

astro-ph.CO

SIDE-real: Supernova Ia Dust Extinction with truncated marginal neural ratio estimation applied to real data

We present the first fully simulation-based hierarchical analysis of the light curves of a population of low-redshift type Ia supernovae (SNae Ia). Our hardware-accelerated forward model, released in the Python package slicsim, includes stochastic variations of each SN's spectral flux distribution (based on the pre-trained BayeSN model), extinction from dust in the host and in the Milky Way, redshift, and realistic instrumental noise. By utilising truncated marginal neural ratio estimation (TMNRE), a neural network-enabled simulation-based inference technique, we implicitly marginalise over 4000 latent variables (for a set of $\approx 100$ SNae Ia) to efficiently infer SN Ia absolute magnitudes and host-galaxy dust properties at the population level while also constraining the parameters of individual objects. Amortisation of the inference procedure allows us to obtain coverage guarantees for our results through Bayesian validation and frequentist calibration. Furthermore, we show a detailed comparison to full likelihood-based inference, implemented through Hamiltonian Monte Carlo, on simulated data and then apply TMNRE to the light curves of 86 SNae Ia from the Carnegie Supernova Project, deriving marginal posteriors in excellent agreement with previous work. Given its ability to accommodate arbitrarily complex extensions to the forward model -- e.g. different populations based on host properties, redshift evolution, complicated photometric redshift estimates, selection effects, and non-Ia contamination -- without significant modifications to the inference procedure, TMNRE has the potential to become the tool of choice for cosmological parameter inference from future, large SN Ia samples.

astro-ph.CO

Scalable hierarchical BayeSN inference: Investigating dependence of SN Ia host galaxy dust properties on stellar mass and redshift

We apply the hierarchical probabilistic SED model BayeSN to analyse a sample of 475 SNe Ia (0.015 < z < 0.4) from Foundation, DES3YR and PS1MD to investigate the properties of dust in their host galaxies. We jointly infer the dust law $R_V$ population distributions at the SED level in high- and low-mass galaxies simultaneously with dust-independent, intrinsic differences. We find an intrinsic mass step of $-0.049\pm0.016$ mag, at a significance of 3.1$σ$, when allowing for a constant intrinsic, achromatic magnitude offset. We additionally apply a model allowing for time- and wavelength-dependent intrinsic differences between SNe Ia in different mass bins, finding $\sim$2$σ$ differences in magnitude and colour around peak and 4.5$σ$ differences at later times. These intrinsic differences are inferred simultaneously with a difference in population mean $R_V$ of $\sim$2$σ$ significance, demonstrating that both intrinsic and extrinsic differences may play a role in causing the host galaxy mass step. We also consider a model which allows the mean of the $R_V$ distribution to linearly evolve with redshift but find no evidence for any evolution - we infer the gradient of this relation $η_R = -0.38\pm0.70$. In addition, we discuss in brief a new, GPU-accelerated Python implementation of BayeSN suitable for application to large surveys which is publicly available and can be used for future cosmological analyses; this code can be found here: https://github.com/bayesn/bayesn.

astro-ph.CO