SearcharxivSearch

arXiv subjects

Derek Bingham

Publications and source records attributed to Derek Bingham.

At least 19 recordsLinked to original sources

Non-Parametric Model Calibration with Stochastic Control Parameters

We present a method for calibrating a computer model using non-parametric techniques where the inputs are stochastic but include calibration parameters whose distributions are unknown and control parameters whose distributions are specified. Our solution gives a distributional estimate over the input space that is consistent with observed field data, while also preserving the distribution of the known marginal of the control parameters. This property is desirable since stochastic inputs often include physical processes affecting the experimental conditions, and a scientifically plausible calibration estimate should preserve well-established distributional properties of these inputs. The method builds on recently developed non-parametric computer model calibration techniques based on the disintegration of measure and Bayesian inference.

stat.ME

Debiasing the Observed Fast Radio Burst Population with the CHIME/FRB Selection Function

The recent release of CHIME/FRB Catalog~2 provides the largest sample to date with which to investigate the intrinsic distributions of fast radio bursts (FRBs). Leveraging an expanded campaign of 587,367 synethetic bursts injected into the live CHIME/FRB search pipeline, we perform a population analysis of the fluence, scattering timescale, pulse width, and dispersion measure distributions of Catalog~2 FRBs. We first infer the intrinsic population using a resampling-based framework that accounts for instrumental selection effects following previous CHIME/FRB population studies. A central goal of this work is to constrain the intrinsic distribution of scattering timescales, that remained weakly constrained in Catalog~1 owing to limited statistics at moderate and large scattering times ($\tau \gtrsim 10\,\mathrm{ms}$ at 600~MHz) and sparse injection coverage in this regime. Second, we construct an explicit multidimensional selection function by training a logistic regression model on the injected events. This model estimates the detection probability as a function of FRB observable properties, including higher-order interaction terms. We incorporate this selection function into a simulation-based inference framework to refine the inferred intrinsic scattering-timescale distribution. We find evidence for a slight downturn in the intrinsic FRB scattering timescale distribution, though a flat or slightly rising distribution cannot be ruled out, that is further supported through a comparison with the higher-frequency scattering timescale distribution observed by Commensal Real-time ASKAP Fast Transients (CRAFT) survey.

astro-ph.HE

Discovery of 30 Repeating Fast Radio Burst Sources and Uniform Population Statistics of 80 Repeating Sources from CHIME/FRB

We present 30 newly discovered repeating fast radio burst (FRB) sources from the second catalog of bursts detected by the FRB backend on the Canadian Hydrogen Intensity Mapping Experiment (CHIME/FRB). These repeaters have extragalactic dispersion measures (DMs) spanning $99.4-1446.0\ \text{pc cm}^{-3}$ and burst rates between $10^{-5.7}$ and $10^{-0.5}$ hr$^{-1}$ scaled to a fluence threshold of 5 Jy ms. We report evidence of monotonic, linear DM variations in four repeaters on years-long timescales. The newly discovered sources bring CHIME/FRB's total number of observed repeating FRBs to 80, 79 of which were discovered by CHIME/FRB, between 2018 July 25 and 2023 September 15. In the full CHIME/FRB sample, only 2.4$\pm 0.4\%$ of sources have been observed to repeat, and we do not find evidence for significant evolution of this value over the duration of the experiment. We find no substantial evidence for bimodal populations of one-off and repeating FRBs in their burst rate distributions; the distribution of upper limits on repeat rates implied from observations of as-yet one-offs is entirely contained within the observed range of repeater burst rates and the distributions do not appear inconsistent. Similarly, using the population analysis framework of C. W. James (2023), we find that our observations of repeating and yet-one-off FRBs are equally well fit assuming a power-law distribution of repeat rates with 50$-$100% of the population repeating.

astro-ph.HE

GPU-accelerated Bayesian inference for block-cave geometry recovery via muon tomography

We describe a Bayesian framework for the inverse problem of geometry recovery of block caving via muon tomography. We work with a low dimensional surface-based representation of the geometry of the block cave, which dramatically reduces the computational requirements of the model while allowing realistic geometries. Adopting a Bayesian approach, we define a prior distribution on the space of geometries that favors realistic cave shapes. Pairing this prior with a likelihood based on the muon tomography forward model, we obtain a posterior distribution over cave geometries using Bayes rule. We obtain approximate samples from this posterior distribution using Markov chain Monte Carlo algorithms running on GPUs, resulting in fast and accurate sampling. We test the fidelity of our methodology by applying it to a simulated block caving scenario for which the ground truth is known. Results show that our method produces sensible geometries that are simultaneously compatible with the data.

stat.AP

Continuity of the Solution of a Non-Parametric Bayesian Statistical Calibration Procedure

Recent work has developed a non-parametric Bayesian approach to the calibration of a computer model, which abstractly amounts to the inversion of a pushforward of stochastic input parameters by a smooth map. The framework has been used in several complex scientific applications, motivating our investigation on the continuity of the solution operator with respect to the distribution on the input parameters. We demonstrate that the solution operator for this approach is uniformly continuous in the total variation metric and weakly continuous for a broad class of distributions.

stat.ME

Nonparametric Bayesian Calibration of Computer Models

Combining field data and computer models is a crucial step for making inferences, predictions, and decisions for complex science and engineering systems. We formulate and analyze a nonparametric Bayesian methodology for calibrating the distribution of parameters in a computer model using field observations. Our results include establishing; a unique nonparametric Bayesian posterior corresponding to a chosen prior with an explicit formula for the posterior density; a maximum entropy property of the posterior corresponding to the uniform prior; the almost everywhere continuity of the posterior density; and a comprehensive statistical analysis of an estimator based on importance sampling. They also include establishing the well-posedness of the nonparametric Bayesian solution of the calibration problem. We illustrate the results using several examples.

stat.ME

Deep Gaussian Process Emulation and Uncertainty Quantification for Large Computer Experiments

Computer models are used as a way to explore complex physical systems. Stationary Gaussian process emulators, with their accompanying uncertainty quantification, are popular surrogates for computer models. However, many computer models are not well represented by stationary Gaussian processes models. Deep Gaussian processes have been shown to be capable of capturing non-stationary behaviors and abrupt regime changes in the computer model response. In this paper, we explore the properties of two deep Gaussian process formulations within the context of computer model emulation. For one of these formulations, we introduce a new parameter that controls the amount of smoothness in the deep Gaussian process layers. We adapt a stochastic variational approach to inference for this model, allowing for prior specification and posterior exploration of the smoothness of the response surface. Our approach can be applied to a large class of computer models, and scales to arbitrarily large simulation designs. The proposed methodology was motivated by the need to emulate an astrophysical model of the formation of binary black hole mergers.

stat.ME

Enhancing Approximate Modular Bayesian Inference by Emulating the Conditional Posterior

In modular Bayesian analyses, complex models are composed of distinct modules, each representing different aspects of the data or prior information. In this context, fully Bayesian approaches can sometimes lead to undesirable feedback between modules, compromising the integrity of the inference. This paper focuses on the "cut-distribution" which prevents unwanted influence between modules by "cutting" feedback. The multiple imputation (DS) algorithm is standard practice for approximating the cut-distribution, but it can be computationally intensive, especially when the number of imputations required is large. An enhanced method is proposed, the Emulating the Conditional Posterior (ECP) algorithm, which leverages emulation to increase the number of imputations. Through numerical experiment it is demonstrated that the ECP algorithm outperforms the traditional DS approach in terms of accuracy and computational efficiency, particularly when resources are constrained. It is also shown how the DS algorithm can be improved using ideas from design of experiments. This work also provides practical recommendations on algorithm choice based on the computational demands of sampling from the prior and cut-distributions.

stat.ME

Rare Event Classification with Weighted Logistic Regression for Identifying Repeating Fast Radio Bursts

An important task in the study of fast radio bursts (FRBs) remains the automatic classification of repeating and non-repeating sources based on their morphological properties. We propose a statistical model that considers a modified logistic regression to classify FRB sources. The classical logistic regression model is modified to accommodate the small proportion of repeaters in the data, a feature that is likely due to the sampling procedure and duration and is not a characteristic of the population of FRB sources. The weighted logistic regression hinges on the choice of a tuning parameter that represents the true proportion $\tau$ of repeating FRB sources in the entire population. The proposed method has a sound statistical foundation, direct interpretability, and operates with only 5 parameters, enabling quicker retraining with added data. Using the CHIME/FRB Collaboration sample of repeating and non-repeating FRBs and numerical experiments, we achieve a classification accuracy for repeaters of nearly 75\% or higher when $\tau$ is set in the range of $50$ to $60$\%. This implies a tentative high proportion of repeaters, which is surprising, but is also in agreement with recent estimates of $\tau$ that are obtained using other methods.

astro-ph.HE

K-Contact Distance for Noisy Nonhomogeneous Spatial Point Data with application to Repeating Fast Radio Burst sources

This paper introduces an approach to analyze nonhomogeneous Poisson processes (NHPP) observed with noise, focusing on previously unstudied second-order characteristics of the noisy process. Utilizing a hierarchical Bayesian model with noisy data, we estimate hyperparameters governing a physically motivated NHPP intensity. Simulation studies demonstrate the reliability of this methodology in accurately estimating hyperparameters. Leveraging the posterior distribution, we then infer the probability of detecting a certain number of events within a given radius, the $k$-contact distance. We demonstrate our methodology with an application to observations of fast radio bursts (FRBs) detected by the Canadian Hydrogen Intensity Mapping Experiment's FRB Project (CHIME/FRB). This approach allows us to identify repeating FRB sources by bounding or directly simulating the probability of observing $k$ physically independent sources within some radius in the detection domain, or the $\textit{probability of coincidence}$ ($P_{\text{C}}$). The new methodology improves the repeater detection $P_{\text{C}}$ in 91% of cases when applied to the largest sample of previously classified observations, with a median improvement factor (existing metric over $P_{\text{C}}$ from our methodology) of $\sim$ 4800.

stat.AP

Fast Emulation, Modular Calibration, and Active Learning for Simulators with Functional Response

Scalable surrogate models enable efficient emulation of computer models (or simulators), particularly when dealing with large ensembles of runs. While Gaussian process (GP) models are commonly employed for emulation, they face limitations in scaling to large datasets. Furthermore, when dealing with dense functional output, such as spatial or time-series data, additional complexities arise, requiring careful handling to ensure fast emulation. This work presents a highly scalable emulator for functional data incorporating local Gaussian process regression. The emulator utilizes global GP lengthscale parameter estimates to scale the input space, leading to a substantial improvement in prediction speed. We demonstrate that our fast approximation-based emulator can serve as a viable alternative to a fully Bayesian approach for functional response, while drastically reducing computational costs. The proposed emulator is applied to quickly calibrate a multiphysics continuum hydrodynamics simulator with a large ensemble of 20000 runs. The methods presented are implemented in the R package FlaGP.

stat.ME

The Mira-Titan Universe IV. High Precision Power Spectrum Emulation

Modern cosmological surveys are delivering datasets characterized by unprecedented quality and statistical completeness; this trend is expected to continue into the future as new ground- and space-based surveys come online. In order to maximally extract cosmological information from these observations, matching theoretical predictions are needed. At low redshifts, the surveys probe the nonlinear regime of structure formation where cosmological simulations are the primary means of obtaining the required information. The computational cost of sufficiently resolved large-volume simulations makes it prohibitive to run very large ensembles. Nevertheless, precision emulators built on a tractable number of high-quality simulations can be used to build very fast prediction schemes to enable a variety of cosmological inference studies. We have recently introduced the Mira-Titan Universe simulation suite designed to construct emulators for a range of cosmological probes. The suite covers the standard six cosmological parameters $\{\omega_m,\omega_b, \sigma_8, h, n_s, w_0\}$ and, in addition, includes massive neutrinos and a dynamical dark energy equation of state, $\{\omega_{\nu}, w_a\}$. In this paper we present the final emulator for the matter power spectrum based on 111 cosmological simulations, each covering a (2.1Gpc)$^3$ volume and evolving 3200$^3$ particles. An additional set of 1776 lower-resolution simulations and TimeRG perturbation theory results for the power spectrum are used to cover scales straddling the linear to mildly nonlinear regimes. The emulator provides predictions at the two to three percent level of accuracy over a wide range of cosmological parameters and is publicly released as part of this paper.

astro-ph.CO

Let's practice what we preach: Planning and interpreting simulation studies with design and analysis of experiments

Statisticians recommend the Design and Analysis of Experiments (DAE) for evidence-based research but often use tables to present their own simulation studies. Could DAE do better? We outline how DAE methods can be used to plan and analyze simulation studies. Tools for planning include fishbone diagrams, factorial and fractional factorial designs. Analysis is carried out via ANOVA, main-effect and interaction plots and other DAE tools. We also demonstrate how Taguchi Robust Parameter Design can be used to study the robustness of methods to a variety of uncontrollable population parameters.

stat.ME

Uncertainty Quantification of a Computer Model for Binary Black Hole Formation

In this paper, a fast and parallelizable method based on Gaussian Processes (GPs) is introduced to emulate computer models that simulate the formation of binary black holes (BBHs) through the evolution of pairs of massive stars. Two obstacles that arise in this application are the a priori unknown conditions of BBH formation and the large scale of the simulation data. We address them by proposing a local emulator which combines a GP classifier and a GP regression model. The resulting emulator can also be utilized in planning future computer simulations through a proposed criterion for sequential design. By propagating uncertainties of simulation input through the emulator, we are able to obtain the distribution of BBH properties under the distribution of physical parameters.

astro-ph.IM

LRP2020: Astrostatistics in Canada

(Abridged from Executive Summary) This white paper focuses on the interdisciplinary fields of astrostatistics and astroinformatics, in which modern statistical and computational methods are applied to and developed for astronomical data. Astrostatistics and astroinformatics have grown dramatically in the past ten years, with international organizations, societies, conferences, workshops, and summer schools becoming the norm. Canada's formal role in astrostatistics and astroinformatics has been relatively limited, but there is a great opportunity and necessity for growth in this area. We conducted a survey of astronomers in Canada to gain information on the training mechanisms through which we learn statistical methods and to identify areas for improvement. In general, the results of our survey indicate that while astronomers see statistical methods as critically important for their research, they lack focused training in this area and wish they had received more formal training during all stages of education and professional development. These findings inform our recommendations for the LRP2020 on how to increase interdisciplinary connections between astronomy and statistics at the institutional, national, and international levels over the next ten years. We recommend specific, actionable ways to increase these connections, and discuss how interdisciplinary work can benefit not only research but also astronomy's role in training Highly Qualified Personnel (HQP) in Canada.

astro-ph.IM

The Mira-Titan Universe II: Matter Power Spectrum Emulation

We introduce a new cosmic emulator for the matter power spectrum covering eight cosmological parameters. Targeted at optical surveys, the emulator provides accurate predictions out to a wavenumber k~5/Mpc and redshift z<=2. Besides covering the standard set of LCDM parameters, massive neutrinos and a dynamical dark energy of state are included. The emulator is built on a sample set of 36 cosmological models, carefully chosen to provide accurate predictions over the wide and large parameter space. For each model, we have performed a high-resolution simulation, augmented with sixteen medium-resolution simulations and TimeRG perturbation theory results to provide accurate coverage of a wide k-range; the dataset generated as part of this project is more than 1.2Pbyte. With the current set of simulated models, we achieve an accuracy of approximately 4%. Because the sampling approach used here has established convergence and error-control properties, follow-on results with more than a hundred cosmological models will soon achieve ~1% accuracy. We compare our approach with other prediction schemes that are based on halo model ideas and remapping approaches. The new emulator code is publicly available.

astro-ph.CO

Estimating parameter uncertainty in binding-energy models by the frequency-domain bootstrap

We propose using the frequency-domain bootstrap (FDB) to estimate errors of modeling parameters when the modeling error is itself a major source of uncertainty. Unlike the usual bootstrap or the simple $χ^2$ analysis, the FDB can take into account correlations between errors. It is also very fast compared to the the Gaussian process Bayesian estimate as often implemented for computer model calibration. The method is illustrated drop model of nuclear binding energies. We find that the FDB gives a more conservative estimate of the uncertainty in liquid drop parameters in better accord with more empirical estimates. For the nuclear physics application, there no apparent obstacle to apply the method to the more accurate and detailed models based on density-functional theory.

nucl-th

A regional compound Poisson process for hurricane and tropical storm damage

In light of intense hurricane activity along the U.S. Atlantic coast, attention has turned to understanding both the economic impact and behaviour of these storms. The compound Poisson-lognormal process has been proposed as a model for aggregate storm damage, but does not shed light on regional analysis since storm path data are not used. In this paper, we propose a fully Bayesian regional prediction model which uses conditional autoregressive (CAR) models to account for both storm paths and spatial patterns for storm damage. When fitted to historical data, the analysis from our model both confirms previous findings and reveals new insights on regional storm tendencies. Posterior predictive samples can also be used for pricing regional insurance premiums, which we illustrate using three different risk measures.

stat.AP