SearcharxivSearch

arXiv subjects

Louis Lyons

Publications and source records attributed to Louis Lyons.

16 recordsLinked to original sources

VERaiPHY -- Validation & Evaluation for Robust AI in PHYsics

Modern machine learning is leading to substantial gains in precision, flexibility, and computational efficiency in fundamental physics. Statistical validation, uncertainty quantification, and robustness assessment are less systematically addressed. The VERaiPHY initiative (Validation & Evaluation for Robust AI in PHYsics) is a series of articles developed within the PHYSTAT programme, aimed at establishing statistical standards for the development, evaluation, and deployment of ML techniques. Each article focuses on a specific methodological domain from a statistics perspective and clarifies statistical questions, tests, and the interpretation of results. This opening article establishes the probabilistic, statistical, and machine learning foundations that the later contributions assume, together with the notation used throughout.

hep-ph

How to Incorporate Systematic Effects into Parameter Determination

We describe two different approaches for incorporating systematics into analyses for parameter determination in the physical sciences. We refer to these as the Pragmatic and the Full methods, with the latter coming in two variants: Full Likelihood and Fully Bayesian. By the use of a simple and readily understood example, we point out the advantage of using the Full Likelihood and Fully Bayesian approaches; a more realistic example from Astrophysics is also presented. This could be relevant for data analyses in a wide range of scientific fields, for situations where systematic effects need to be incorporated in the analysis procedure. This note is an extension of part of the talk by van Dyk at the PHYSTAT-Systematics meeting.

hep-ex

Statistical techniques to estimate the SARS-CoV-2 infection fatality rate

The determination of the infection fatality rate (IFR) for the novel SARS-CoV-2 coronavirus is a key aim for many of the field studies that are currently being undertaken in response to the pandemic. The IFR together with the basic reproduction number $R_0$, are the main epidemic parameters describing severity and transmissibility of the virus, respectively. The IFR can be also used as a basis for estimating and monitoring the number of infected individuals in a population, which may be subsequently used to inform policy decisions relating to public health interventions and lockdown strategies. The interpretation of IFR measurements requires the calculation of confidence intervals. We present a number of statistical methods that are relevant in this context and develop an inverse problem formulation to determine correction factors to mitigate time-dependent effects that can lead to biased IFR estimates. We also review a number of methods to combine IFR estimates from multiple independent studies, provide example calculations throughout this note and conclude with a summary and "best practice" recommendations. The developed code is available online.

stat.AP

Reproducibility and Replication of Experimental Particle Physics Results

Recently, much attention has been focused on the replicability of scientific results, causing scientists, statisticians, and journal editors to examine closely their methodologies and publishing criteria. Experimental particle physicists have been aware of the precursors of non-replicable research for many decades and have many safeguards to ensure that the published results are as reliable as possible. The experiments require large investments of time and effort to design, construct, and operate. Large collaborations produce and check the results, and many papers are signed by more than three thousand authors. This paper gives an introduction to what experimental particle physics is and to some of the tools that are used to analyze the data. It describes the procedures used to ensure that results can be computationally reproduced, both by collaborators and by non-collaborators. It describes the status of publicly available data sets and analysis tools that aid in reproduction and recasting of experimental results. It also describes methods particle physicists use to maximize the reliability of the results, which increases the probability that they can be replicated by other collaborations or even the same collaborations with more data and new personnel. Examples of results that were later found to be false are given, both with failed replication attempts and one with alarmingly successful replications. While some of the characteristics of particle physics experiments are unique, many of the procedures and techniques can be and are used in other fields.

physics.data-an

PHYSTAT$\nu$ at CERN (January 2019)

A short overview is provided of the recent PHYSTAT$\nu$ meeting at CERN, which dealt with statistical issues relevant for neutrino experiments.

hep-ex

A Paradox about Likelihood Ratios?

We consider whether the asymptotic distributions for the log-likelihood ratio test statistic are expected to be Gaussian or chi-squared. Two straightforward examples provide insight on the difference.

physics.data-an

Statistical Issues in Searches for New Physics

Given the cost, both financial and even more importantly in terms of human effort, in building High Energy Physics accelerators and detectors and running them, it is important to use good statistical techniques in analysing data. Some of the statistical issues that arise in searches for New Physics are discussed briefly. They include topics such as: Should we insist on the 5 sigma criterion for discovery claims? The probability of A, given B, is not the same as the probability of B, given A. The meaning of p-values. What is Wilks Theorem and when does it not apply? How should we deal with the `Look Elsewhere Effect'? Dealing with systematics such as background parametrisation. Coverage: What is it and does my method have the correct coverage? The use of p0 versus p1 plots.

hep-ex

Testing Hypotheses in Particle Physics: Plots of $p_{0}$ Versus $p_{1}$

For situations where we are trying to decide which of two hypotheses $H_{0}$ and $H_{1}$ provides a better description of some data, we discuss the usefulness of plots of $p_{0}$ versus $p_{1}$, where $p_{i}$ is the $p$-value for testing $H_{i}$. They provide an interesting way of understanding the difference between the standard way of excluding $H_{1}$ and the $CL_{s}$ approach; the Punzi definition of sensitivity; the relationship between $p$-values and likelihood ratios; and the probability of observing misleading evidence. They also help illustrate the Law of the Iterated Logarithm and the Jeffreys-Lindley paradox.

stat.ME

Raster scan or 2-D approach?

We consider the relative merits of two different approaches to discovery or exclusion of new phenomena, a raster scan or a 2-dimensional approach.

hep-ex

Discovering the Significance of 5 sigma

We discuss the traditional criterion for discovery in Particle Physics of requiring a significance corresponding to at least 5 sigma; and whether a more nuanced approach might be better.

physics.data-an

Bayes and Frequentism: a Particle Physicist's perspective

In almost every scientific field, an experiment involves collecting data and then analysing it. The analysis stage will often consist in trying to extract some physical parameter and estimating its uncertainty; this is known as Parameter Determination. An example would be the determination of the mass of the top quark, from data collected from high energy proton-proton collisions. A different aim is to choose between two possible hypotheses. For example, are data on the recession speed s of distant galaxies proportional to their distance d, or do they fit better to a model where the expansion of the Universe is accelerating? There are two fundamental approaches to such statistical analyses - Bayesian and Frequentist. This article discusses the way they differ in their approach to probability, and then goes on to consider how this affects the way they deal with Parameter Determination and Hypothesis Testing. The examples are taken from every-day life and from Particle Physics.

physics.data-an

Open statistical issues in particle physics

Many statistical issues arise in the analysis of Particle Physics experiments. We give a brief introduction to Particle Physics, before describing the techniques used by Particle Physicists for dealing with statistical problems, and also some of the open statistical questions.

stat.AP

Interval estimation in the presence of nuisance parameters. 1. Bayesian approach

We address the common problem of calculating intervals in the presence of systematic uncertainties. We aim to investigate several approaches, but here describe just a Bayesian technique for setting upper limits. The particular example we study is that of inferring the rate of a Poisson process when there are uncertainties on the acceptance and the background. Limit calculating software associated with this work is available in the form of C functions.

physics.data-an