SearcharxivSearch

arXiv · 2501.16462

Optimizing Neural Network Surrogate Models: Application to Black Hole Merger Remnants

Abstract

Surrogate models of numerical relativity simulations of merging black holes provide the most accurate tools for gravitational-wave data analysis. Neural network-based surrogates promise evaluation speedups, but their accuracy relies on (often obscure) tuning of settings such as the network architecture, hyperparameters, and the size of the training dataset. We propose a systematic optimization strategy that formalizes setting choices and motivates the amount of training data required. We apply this strategy on NRSur7dq4Remnant, an existing surrogate model for the properties of the remnant of generically precessing binary black hole mergers and construct a neural network version, which we label NRSur7dq4Remnant_NN. The systematic optimization strategy results in a new surrogate model with comparable accuracy, and provides insights into the meaning and role of the various network settings and hyperparameters as well as the structure of the physical process. Moreover, NRSur7dq4Remnant_NN results in evaluation speedups of up to $8$ times on a single CPU and a further improvement of $2,000$ times when evaluated in batches on a GPU. To determine the training set size, we propose an iterative enrichment strategy that efficiently samples the parameter space using much smaller training sets than naive sampling. NRSur7dq4Remnant_NN requires $O(10^4)$ training data, so neural network-based surrogates are ideal for speeding-up models that support such large training datasets, but at the moment cannot directly be applied to numerical relativity catalogs that are $O(10^3)$ in size. The optimization strategy is available through the gwbonsai package.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lucy M. Thomas, Katerina Chatziioannou, Vijay Varma, Scott E. Field. 2025-01-27. Optimizing Neural Network Surrogate Models: Application to Black Hole Merger Remnants. https://doi.org/10.1103/physrevd.111.104029

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Electrovacuum Black Hole Uniqueness

We prove the black hole uniqueness conjecture in the axially symmetric, stationary, electrovacuum setting, subject to the refined asymptotic analysis of the associated singular harmonic maps, which includes an analyticity hypothesis at the axes. More precisely, it is shown that any asymptotically flat solution of the Einstein--Maxwell equations in this class, with more than one black hole horizon component is either: Majumdar--Papapetrou, up to a duality rotation, in which case all logarithmic angle defects vanish, or every finite axis rod logarithmic angle defect is strictly negative and hence every interaction force is strictly attractive. The proof extends the singular harmonic map method used for vacuum Kerr uniqueness in [18].

gr-qc

Constraining Modified Mass-to-Horizon Cosmology Through Primordial Inflationary Observables

We investigate slow-roll inflation in a modified cosmological framework inspired by a generalized mass-to-horizon relation (MHR), $M=\gamma {c^2 L^n}/{G}$, where $n$ is a real parameter and $\gamma$ a dimensional constant. Using Padmanabhan's emergence paradigm, we derive the modified Friedmann equations for a flat FRW universe and analyze the dynamics of a canonical scalar field (inflaton) under the slow-roll approximation. We study the resulting inflationary phenomenology for power-law and Starobinsky potentials. For power-law potentials, the MHR modification fails to reconcile these models with current CMB constraints on $r$ and $n_s$. In contrast, Starobinsky inflation exhibits significant sensitivity to deviations from $n=1$. A perturbative analysis ($n=1+\Delta$) yields corrections to inflationary observables. We observe that the scalar power-spectrum normalization, under a fixed-Starobinsky prescription, imposes the stringent constraint $0.960 \lesssim n \lesssim 1.040$ for $N=60$ efolds. This is considerably tighter than spectral-index bounds. Our results establish inflation, particularly Starobinsky-like models, as a sensitive probe of generalized horizon thermodynamics and departures from standard MHR scaling.

gr-qc

Improving the Sensitivity of Gravitational Wave Detection with Weighted Conformal Prediction

In the last decade, kilometre-scale interferometric gravitational-wave detectors have observed hundreds of compact binary mergers, the majority of which are binary black holes. However, the data are noise-dominated, and multiple independent search algorithms (pipelines) are used to enhance sensitivity and improve robustness. Rather than the standard approach of selecting the most significant pipeline output, we combine the outputs from all pipelines using a conformal prediction-based framework to provide statistically rigorous confidence estimates for candidate events. While combining pipelines improves sensitivity and ranking robustness, it requires a principled statistical framework that remains valid as data properties evolve across observing runs. A key challenge is distribution shifts between simulated datasets used for training and calibration and the real, unlabelled, observations used for testing, which can invalidate coverage guarantees and bias confidence estimates. In this work, we address this challenge by incorporating likelihood-ratio reweighting into our conformal prediction framework to account for covariate shift. Using mock datasets containing simulated signals, we demonstrate that weighted conformal prediction restores well-calibrated coverage under covariate shift and increases the confidence of events near the detection threshold, recovering true signals that would otherwise be missed.

gr-qc