SearcharxivSearch

arXiv subjects

Sara Algeri

Publications and source records attributed to Sara Algeri.

At least 19 recordsLinked to original sources

Compensator-based inference for signal detection under unknown background: the binned data case

The problem of signal detection under an unknown background can be framed as one of inferring the weight of a mixture model with one misspecified component. Banerjee and Algeri (2026) show that, for this problem, the conservativeness of the inference is entirely determined by one single parameter, called the compensator. They demonstrate that, when the data are independent and identically distributed, an inferential approach based on the compensator circumvents the need to estimate the density of the misspecified component and the associated challenges. The main purpose of this manuscript is to broaden the scope of such an approach and extend it to the case in which, as is often encountered in modern experiments in physics and astronomy, the data consist of Poisson counts observed over a large number of bins.

stat.ME

Compensator-Based Inference for Signal Detection Under Unknown Background

The problem of detecting new signals in the presence of an unknown background is ubiquitous in scientific discoveries and is especially prominent in the physical sciences. Most solutions proposed thus far to address the problem focus on estimating the background distribution and using that estimate to infer the signal. By studying the geometry of the problem, this article demonstrates that estimating the background distribution is somewhat unnecessary for inferring the signal intensity. Instead, it suffices to estimate a single parameter, referred to as the compensator, to account for the incomplete knowledge on the background, substantially simplifying the problem's complexity and enabling proper uncertainty propagation. Such a compensator is shown to govern the conservativeness of the inference, both in the proposed setup and in likelihood-based approaches.

stat.ME

A scalable Bayesian framework for galaxy emission line detection and redshift estimation

Estimating galaxy redshifts is crucial for constraining key physical quantities like those in the equation of state of dark energy. Modern telescopes such as the James Webb Space Telescope, the Euclid Space Telescope, and the NASA Nancy Grace Roman Space Telescope are producing massive amounts of spectroscopic data that enable precise redshift estimation. However, a galaxy's redshift can be estimated only when emission lines are present in the observed spectrum, which is unknown a priori. A novel Bayesian approach to estimating redshift and simultaneously testing for the presence of emission lines is developed. Although modern spectroscopic surveys involve millions of spectra and give rise to highly multimodal posterior distributions, the proposed framework remains computationally efficient, admitting a parallelizable implementation suitable for large-scale inference.

astro-ph.IM

A New Class of Asymptotically Distribution-Free Smooth Tests

This article demonstrates how recent developments in the theory of empirical processes allow us to construct a new family of asymptotically distribution-free smooth tests. Their distribution-free property is preserved even when the parameters are estimated, model selection is performed, and the sample size is only moderately large. A computationally efficient alternative to the classical parametric bootstrap is also discussed.

math.ST

On Validating Angular Power Spectral Models for the Stochastic Gravitational-Wave Background Without Distributional Assumptions

It is demonstrated that estimators of the angular power spectrum commonly used for the stochastic gravitational-wave background (SGWB) lack a closed-form analytical expression for the likelihood function and, typically, cannot be accurately approximated by a Gaussian likelihood. Nevertheless, a robust statistical analysis can be performed to enable the estimation and testing of angular power spectral models for the SGWB without specifying distributional assumptions. Here, the technical aspects of the method are discussed in detail. Moreover, a new, consistent estimator for the covariance of the angular power spectrum is derived. The proposed approach is applied to data from the third observing run (O3) of Advanced LIGO and Advanced Virgo.

astro-ph.IM

Testing models for angular power spectra: A distribution-free approach

A novel goodness-of-fit strategy is introduced for testing models of angular power spectra with unknown parameters. Using this strategy, it is possible to assess the validity of such models without specifying the distribution of the angular power spectrum estimators. This holds under general conditions, ensuring the method's applicability in diverse applications. Moreover, the proposed solution overcomes the need for case-by-case simulations when testing different models, leading to notable computational advantages.

physics.data-an

On the statistical analysis of grouped data: when Pearson $\chi^2$ and other divisible statistics are not goodness-of-fit tests

Thousands of experiments are analyzed, and papers are published each year involving the statistical analysis of grouped data. While this area of statistics is often perceived -- somewhat naively -- as saturated, several misconceptions still affect everyday practice, and new frontiers have so far remained unexplored. Researchers must be aware of the limitations affecting their analyses and what new possibilities are at their hands. The article introduces a unifying approach to the analysis of divisible statistics -- that includes Pearson's $\chi^2$, the likelihood ratio, and spectral statistics, as special cases -- when a statistician deals with a large number of bins/groups, thus leading to a large number of small or moderate frequencies. Performance of the tests is analyzed against the class of contiguous (local) alternatives. Perhaps the most surprising result here is that, in this `sparse' regime, most of the tests proposed in the literature can be modified to produce more powerful tests, and no single test based on a divisible statistic leads to a goodness-of-fit test. Distribution-free goodness-of-fit tests are also constructed.

stat.ME

A novel approach to detect line emission under high background in high-resolution X-ray spectra

We develop a novel statistical approach to identify emission features or set upper limits in high-resolution spectra in the presence of high background. The method relies on detecting differences from the background using smooth tests and using classical likelihood ratio tests to characterise known shapes like emission lines. We perform signal detection or place upper limits on line fluxes while accounting for the problem of multiple comparisons. We illustrate the method by applying it to a Chandra LETGS+HRC-S observation of symbiotic star RT Cru, successfully detecting previously known features like the Fe line emission in the 6-7 keV range and the Iridium-edge due to the mirror coating on Chandra. We search for thermal emission lines from Ne X, Fe XVII, O VIII, and O VII, but do not detect them, and place upper limits on their intensities consistent with a $\approx$1 keV plasma. We serendipitously detect a line at 16.93 $\unicode{x212B}$ that we attribute to photoionisation or a reflection component.

astro-ph.IM

Sequential hypothesis testing for Axion Haloscopes

The goal of this paper is to introduce a novel likelihood-based inferential framework for axion haloscopes which is valid under the commonly applied "rescanning" protocol. The proposed method enjoys short data acquisition times and a simple tuning of the detector configuration. Local statistical significance and power are computed analytically, avoiding the need of burdensome simulations. Adequate corrections for the look-elsewhere effect are also discussed. The performance of our inferential strategy is compared with that of a simple method which exploits the geometric probability of rescan. Finally, we exemplify the method with an application to a HAYSTAC type axion haloscope.

physics.data-an

Estimating the lifetime risk of a false positive screening test result

False positive results in screening tests have potentially severe psychological, medical, and financial consequences for the recipient. However, there have been few efforts to quantify how the risk of a false positive accumulates over time. We seek to fill this gap by estimating the probability that an individual who adheres to the U.S. Preventive Services Task Force (USPSTF) screening guidelines will receive at least one false positive in a lifetime. To do so, we assembled a data set of 116 studies cited by the USPSTF that report the number of true positives, false negatives, true negatives, and false positives for the primary screening procedure for one of five cancers or six sexually transmitted diseases. We use these data to estimate the probability that an individual in one of 14 demographic subpopulations will receive at least one false positive for one of these eleven diseases in a lifetime. We specify a suitable statistical model to account for the hierarchical structure of the data, and we use the parametric bootstrap to quantify the uncertainty surrounding our estimates. The estimated probability of receiving at least one false positive in a lifetime is 85.5% ($\pm$0.9%) and 38.9% ($\pm$3.6%) for baseline groups of women and men, respectively. It is higher for subpopulations recommended to screen more frequently than the baseline, including more vulnerable groups such as pregnant women and men who have sex with men. Since screening technology is imperfect, false positives remain inevitable. The high lifetime risk of a false positive reveals the importance of educating patients about this phenomenon.

stat.AP

Informative Goodness-of-Fit for Multivariate Distributions

This article introduces an informative goodness-of-fit (iGOF) approach to study multivariate distributions. When the null model is rejected, iGOF allows us to identify the underlying sources of mismodeling and naturally equips practitioners with additional insights on the nature of the deviations from the true distribution. The informative character of the procedure is achieved by exploiting smooth tests and random fields theory to facilitate the analysis of multivariate data. Simulation studies show that iGOF enjoys high power for different types of alternatives. The methods presented here directly address the problem of background mismodeling arising in physics and astronomy. It is in these areas that the motivation of this work is rooted.

stat.ME

K-2 rotated goodness-of-fit for multivariate data

Consider a set of multivariate distributions, $F_1,\dots,F_M$, aiming to explain the same phenomenon. For instance, each $F_m$ may correspond to a different candidate background model for calibration data, or to one of many possible signal models we aim to validate on experimental data. In this article, we show that tests for a wide class of apparently different models $F_{m}$ can be mapped into a single test for a reference distribution $Q$. As a result, valid inference for each $F_m$ can be obtained by simulating \underline{only} the distribution of the test statistic under $Q$. Furthermore, $Q$ can be chosen conveniently simple to substantially reduce the computational time.

stat.ME

Exhaustive goodness-of-fit via smoothed inference and graphics

Classical tests of goodness-of-fit aim to validate the conformity of a postulated model to the data under study. Given their inferential nature, they can be considered a crucial step in confirmatory data analysis. In their standard formulation, however, they do not allow exploring how the hypothesized model deviates from the truth nor do they provide any insight into how the rejected model could be improved to better fit the data. The main goal of this work is to establish a comprehensive framework for goodness-of-fit which naturally integrates modeling, estimation, inference, and graphics. Modeling and estimation focus on a novel formulation of smooth tests that easily extends to arbitrary distributions, either continuous or discrete. Inference and adequate post-selection adjustments are performed via a specially designed smoothed bootstrap and the results are summarized via an exhaustive graphical tool called CD-plot.

stat.ME

Detecting new signals under background mismodelling

Searches for new astrophysical phenomena often involve several sources of non-random uncertainties which can lead to highly misleading results. Among these, model-uncertainty arising from background mismodelling can dramatically compromise the sensitivity of the experiment under study. Specifically, overestimating the background distribution in the signal region increases the chances of missing new physics. Conversely, underestimating the background outside the signal region leads to an artificially enhanced sensitivity and a higher likelihood of claiming false discoveries. The aim of this work is to provide a unified statistical strategy to perform modelling, estimation, inference, and signal characterization under background mismodelling. The method proposed allows to incorporate the (partial) scientific knowledge available on the background distribution and provides a data-updated version of it in a purely nonparametric fashion without requiring the specification of prior distributions on the parameters. Applications in the context of dark matter searches and radio surveys show how the tools presented in this article can be used to incorporate non-stochastic uncertainty due to instrumental noise and to overcome violations of classical distributional assumptions in stacking experiments.

physics.data-an

Searching for new physics with profile likelihoods: Wilks and beyond

Particle physics experiments use likelihood ratio tests extensively to compare hypotheses and to construct confidence intervals. Often, the null distribution of the likelihood ratio test statistic is approximated by a $χ^2$ distribution, following a theorem due to Wilks. However, many circumstances relevant to modern experiments can cause this theorem to fail. In this paper, we review how to identify these situations and construct valid inference.

physics.data-an

Testing One Hypothesis Multiple Times: The Multidimensional Case

The identification of new rare signals in data, the detection of a sudden change in a trend, and the selection of competing models, are among the most challenging problems in statistical practice. These challenges can be tackled using a test of hypothesis where a nuisance parameter is present only under the alternative, and a computationally efficient solution can be obtained by the "Testing One Hypothesis Multiple times" (TOHM) method. In the one-dimensional setting, a fine discretization of the space of the non-identifiable parameter is specified, and a global p-value is obtained by approximating the distribution of the supremum of the resulting stochastic process. In this paper, we propose a computationally efficient inferential tool to perform TOHM in the multidimensional setting. Here, the approximations of interest typically involve the expected Euler Characteristics (EC) of the excursion set of the underlying random field. We introduce a simple algorithm to compute the EC in multiple dimensions and for arbitrary large significance levels. This leads to an highly generalizable computational tool to perform inference under non-standard regularity conditions.

stat.ME

Testing One Hypothesis Multiple times

In applied settings, tests of hypothesis where a nuisance parameter is only identifiable under the alternative often reduces into one of Testing One Hypothesis Multiple times (TOHM). Specifically, a fine discretization of the space of the non-identifiable parameter is specified, and the null hypothesis is tested against a set of sub-alternative hypothesis, one for each point of the discretization. The resulting sub-test statistics are then combined to obtain a global p-value. In this paper, we discuss a computationally efficient inferential tool to perform TOHM under stringent significance requirements, such as those typically required in the physical sciences, (e.g., p-value $<10^{-7}$). The resulting procedure leads to a generalized approach to perform inference under non-standard conditions, including non-nested models comparisons.

stat.ME

Statistical challenges in the search for dark matter

The search for the particle nature of dark matter has given rise to a number of experimental, theoretical and statistical challenges. Here, we report on a number of these statistical challenges and new techniques to address them, as discussed in the DMStat workshop held Feb 26 - Mar 3 2018 at the Banff International Research Station for Mathematical Innovation and Discovery (BIRS) in Banff, Alberta.

hep-ph