Searcharxiv⌕ Search

arXiv subjects

Robert D. Cousins

Publications and source records attributed to Robert D. Cousins.

16 recordsLinked to original sources

Comment on Frank Porter, "Confidence intervals for the Poisson distribution"

Frank Porter has recently posted a review of "Confidence intervals for the Poisson distribution" (arXiv:2509.02852). The long, diverse history of such intervals is closely related to that of confidence intervals for the parameter of the binomial distribution. While much of Porter's paper is enlightening and food for thought, I believe that his discussion of the intervals advocated by Gary Feldman and myself (FC) based on the likelihood ratio test (arXiv:physics/9711021) is flawed. The fundamental point of disagreement is whether or not the likelihood function exists in a part of parameter space where the statistical model does not exist. (The new paper says yes; FC say no.) Here, I focus mainly on that issue and the consequences, along with a few other remarks.

physics.data-an↗

Lectures on Statistics in Theory: Prelude to Statistics in Practice

This is a writeup of lectures on "statistics" that have evolved from the initial version for the 2009 Hadron Collider Physics Summer School at CERN to versions for other venues and, most recently, for the African School of Fundamental Physics and Applications in 2024. The emphasis is on foundations, using simple examples to illustrate the points that are still debated in the professional statistics literature. The three main approaches to interval estimation (Neyman confidence, Bayesian, likelihood ratio) are discussed and compared in detail, with and without nuisance parameters. Hypothesis testing is discussed mainly from the frequentist point of view, with pointers to the Bayesian literature. Various foundational issues are emphasized, including the conditionality principle and the likelihood principle.

physics.data-an↗

PHYSTAT Informal Review: Marginalizing versus Profiling of Nuisance Parameters

This is a writeup, with some elaboration, of the talks by the two authors (a physicist and a statistician) at the first PHYSTAT Informal review on January 24, 2024. We discuss Bayesian and frequentist approaches to dealing with nuisance parameters, in particular, integrated versus profiled likelihood methods. In regular models, with finitely many parameters and large sample sizes, the two approaches are asymptotically equivalent. But, outside this setting, the two methods can lead to different tests and confidence intervals. Assessing which approach is better generally requires comparing the power of the tests or the length of the confidence intervals. This analysis has to be conducted on a case-by-case basis. In the extreme case where the number of nuisance parameters is very large, possibly infinite, neither approach may be useful. Part I provides an informal history of usage in high energy particle physics, including a simple illustrative example. Part II includes an overview of some more recently developed methods in the statistics literature, including methods applicable when the use of the likelihood function is problematic.

physics.data-an↗

Connections between statistical practice in elementary particle physics and the severity concept as discussed in Mayo's Statistical Inference as Severe Testing

For many years, philosopher-of-statistics Deborah Mayo has been advocating the concept of severe testing as a key part of hypothesis testing. Her recent book, Statistical Inference as Severe Testing, is a comprehensive exposition of her arguments in the context of a historical study of many threads of statistical inference, both frequentist and Bayesian. Her foundational point of view is called error statistics, emphasizing frequentist evaluation of the errors called Type I and Type II in the Neyman-Pearson theory of frequentist hypothesis testing. Since the field of elementary particle physics (also known as high energy physics) has strong traditions in frequentist inference, one might expect that something like the severity concept was independently developed in the field. Indeed, I find that, at least operationally (numerically), we high-energy physicists have long interpreted data in ways that map directly onto severity. Whether or not we subscribe to Mayo's philosophical interpretations of severity is a more complicated story that I do not address here.

stat.OT↗

What is the likelihood function, and how is it used in particle physics?

Likelihood functions are ubiquitous in data analyses at the LHC and elsewhere in particle physics. Partly because "probability" and "likelihood" are virtual synonyms in everyday English, but crucially distinct in data analysis, there is great potential for confusion. Furthermore, each of various approaches to statistical inference (likelihoodist, Neyman-Pearson, Bayesian) uses the likelihood function in different ways. This note is intended to provide a brief introduction at the advanced undergraduate or beginning graduate student level, citing a few papers giving examples and containing numerous pointers to the vast literature on likelihood. The Likelihood Principle (routinely violated in particle physics analyses) is mentioned as an unresolved issue in the philosophical foundations of statistics.

physics.data-an↗

Should unfolded histograms be used to test hypotheses?

In many analyses in high energy physics, attempts are made to remove the effects of detector smearing in data by techniques referred to as "unfolding" histograms, thus obtaining estimates of the true values of histogram bin contents. Such unfolded histograms are then compared to theoretical predictions, either to judge the goodness of fit of a theory, or to compare the abilities of two or more theories to describe the data. When doing this, even informally, one is testing hypotheses. However, a more fundamentally sound way to test hypotheses is to smear the theoretical predictions by simulating detector response and then comparing to the data without unfolding; this is also frequently done in high energy physics, particularly in searches for new physics. One can thus ask: to what extent does hypothesis testing after unfolding data materially reproduce the results obtained from testing by smearing theoretical predictions? We argue that this "bottom-line-test" of unfolding methods should be studied more commonly, in addition to common practices of examining variance and bias of estimates of the true contents of histogram bins. We illustrate bottom-line-tests in a simple toy problem with two hypotheses.

physics.data-an↗

The Jeffreys-Lindley Paradox and Discovery Criteria in High Energy Physics

The Jeffreys-Lindley paradox displays how the use of a p-value (or number of standard deviations z) in a frequentist hypothesis test can lead to an inference that is radically different from that of a Bayesian hypothesis test in the form advocated by Harold Jeffreys in the 1930s and common today. The setting is the test of a well-specified null hypothesis (such as the Standard Model of elementary particle physics, possibly with "nuisance parameters") versus a composite alternative (such as the Standard Model plus a new force of nature of unknown strength). The p-value, as well as the ratio of the likelihood under the null hypothesis to the maximized likelihood under the alternative, can strongly disfavor the null hypothesis, while the Bayesian posterior probability for the null hypothesis can be arbitrarily large. The academic statistics literature contains many impassioned comments on this paradox, yet there is no consensus either on its relevance to scientific communication or on its correct resolution. The paradox is quite relevant to frontier research in high energy physics. This paper is an attempt to explain the situation to both physicists and statisticians, in the hope that further progress can be made.

stat.ME↗

Negatively Biased Relevant Subsets Induced by the Most-Powerful One-Sided Upper Confidence Limits for a Bounded Physical Parameter

Suppose an observable x is the measured value (negative or non-negative) of a true mean mu (physically non-negative) in an experiment with a Gaussian resolution function with known fixed rms deviation s. The most powerful one-sided upper confidence limit at 95% C.L. is UL = x+1.64s, which I refer to as the "original diagonal line". Perceived problems in HEP with small or non-physical upper limits for x<0 historically led, for example, to substitution of max(0,x) for x, and eventually to abandonment in the Particle Data Group's Review of Particle Physics of this diagonal line relationship between UL and x. Recently Cowan, Cranmer, Gross, and Vitells (CCGV) have advocated a concept of "power constraint" that when applied to this problem yields variants of diagonal line, including UL = max(-1,x)+1.64s. Thus it is timely to consider again what is problematic about the original diagonal line, and whether or not modifications cure these defects. In a 2002 Comment, statistician Leon Jay Gleser pointed to the literature on recognizable and relevant subsets. For upper limits given by the original diagonal line, the sample space for x has recognizable relevant subsets in which the quoted 95% C.L. is known to be negatively biased (anti-conservative) by a finite amount for all values of mu. This issue is at the heart of a dispute between Jerzy Neyman and Sir Ronald Fisher over fifty years ago, the crux of which is the relevance of pre-data coverage probabilities when making post-data inferences. The literature describes illuminating connections to Bayesian statistics as well. Methods such as that advocated by CCGV have 100% unconditional coverage for certain values of mu and hence formally evade the traditional criteria for negatively biased relevant subsets; I argue that concerns remain. Comparison with frequentist intervals advocated by Feldman and Cousins also sheds light on the issues.

physics.data-an↗

Frequentist Evaluation of Intervals Estimated for a Binomial Parameter and for the Ratio of Poisson Means

Confidence intervals for a binomial parameter or for the ratio of Poisson means are commonly desired in high energy physics (HEP) applications such as measuring a detection efficiency or branching ratio. Due to the discreteness of the data, in both of these problems the frequentist coverage probability unfortunately depends on the unknown parameter. Trade-offs among desiderata have led to numerous sets of intervals in the statistics literature, while in HEP one typically encounters only the classic intervals of Clopper-Pearson (central intervals with no undercoverage but substantial over-coverage) or a few approximate methods which perform rather poorly. If strict coverage is relaxed, some sort of averaging is needed to compare intervals. In most of the statistics literature, this averaging is over different values of the unknown parameter, which is conceptually problematic from the frequentist point of view in which the unknown parameter is typically fixed. In contrast, we perform an (unconditional) {\it average over observed data} in the ratio-of-Poisson-means problem. If strict conditional coverage is desired, we recommend Clopper-Pearson intervals and intervals from inverting the likelihood ratio test (for central and non-central intervals, respectively). Lancaster's mid-$P$ modification to either provides excellent unconditional average coverage in the ratio-of-Poisson-means problem.

physics.data-an↗

Comment on "Bayesian Analysis of Pentaquark Signals from CLAS Data", with Response to the Reply by Ireland and Protopopsecu

The CLAS Collaboration has published an analysis using Bayesian model selection. My Comment criticizing their use of arbitrary prior probability density functions, and a Reply by D.G. Ireland and D. Protopopsecu, have now been published as well. This paper responds to the Reply and discusses the issues in more detail, with particular emphasis on the problems of priors in Bayesian model selection.

hep-ph↗

Annotated Bibliography of Some Papers on Combining Significances or p-values

A question that comes up repeatedly is how to combine the results of two experiments if all that is known is that one experiment had a n-sigma effect and another experiment had a m-sigma effect. This question is not well-posed: depending on what additional assumptions are made, the preferred answer is different. The note lists some of the more prominent papers on the topic, with some brief comments and excerpts.

physics.data-an↗

Evaluation of three methods for calculating statistical significance when incorporating a systematic uncertainty into a test of the background-only hypothesis for a Poisson process

Hypothesis tests for the presence of new sources of Poisson counts amidst background processes are frequently performed in high energy physics (HEP), gamma ray astronomy (GRA), and other branches of science. While there are conceptual issues already when the mean rate of background is precisely known, the issues are even more difficult when the mean background rate has non-negligible uncertainty. After describing a variety of methods to be found in the HEP and GRA literature, we consider in detail three classes of algorithms and evaluate them over a wide range of parameter space, by the criterion of how close the ensemble-average Type I error rate (rejection of the background-only hypothesis when it is true) compares with the nominal significance level given by the algorithm. We recommend wider use of an algorithm firmly grounded in frequentist tests of the ratio of Poisson means, although for very low counts the over-coverage can be severe due to the effect of discreteness. We extend the studies of Cranmer, who found that a popular Bayesian-frequentist hybrid can undercover severely when taken to high Z values. We also examine the profile likelihood method, which has long been used in GRA and HEP; it provides an excellent approximation in much of the parameter space, as previously studied by Rolke and collaborators.

physics.data-an↗

Application of Conditioning to the Gaussian-with-Boundary Problem in the Unified Approach to Confidence Intervals

Roe and Woodroofe (RW) have suggested that certain conditional probabilities be incorporated into the ``unified approach'' for constructing confidence intervals, previously described by Feldman and Cousins (FC). RW illustrated this conditioning technique using one of the two prototype problems in the FC paper, that of Poisson processes with background. The main effect was on the upper curve in the confidence belt. In this paper, we attempt to apply this style of conditioning to the other prototype problem, that of Gaussian errors with a bounded physical region. We find that the lower curve on the confidence belt is also moved significantly, in an undesirable manner.

physics.data-an↗

Kalman Filter Track Fits and Track Breakpoint Analysis

We give an overview of track fitting using the Kalman filter method in the NOMAD detector at CERN, and emphasize how the wealth of by-product information can be used to analyze track breakpoints (discontinuities in track parameters caused by scattering, decay, etc.). After reviewing how this information has been previously exploited by others, we describe extensions which add power to breakpoint detection and characterization. We show how complete fits to the entire track, with breakpoint parameters added, can be easily obtained from the information from unbroken fits. Tests inspired by the Fisher F-test can then be used to judge breakpoints. Signed quantities (such as change in momentum at the breakpoint) can supplement unsigned quantities such as the various chisquares. We illustrate the method with electrons from real data, and with Monte Carlo simulations of pion decays.

physics.data-an↗

A Unified Approach to the Classical Statistical Analysis of Small Signals

We give a classical confidence belt construction which unifies the treatment of upper confidence limits for null results and two-sided confidence intervals for non-null results. The unified treatment solves a problem (apparently not previously recognized) that the choice of upper limit or two-sided intervals leads to intervals which are not confidence intervals if the choice is based on the data. We apply the construction to two related problems which have recently been a battle-ground between classical and Bayesian statistics: Poisson processes with background, and Gaussian errors with a bounded physical region. In contrast with the usual classical construction for upper limits, our construction avoids unphysical confidence intervals. In contrast with some popular Bayesian intervals, our intervals eliminate conservatism (frequentist coverage greater than the stated confidence) in the Gaussian case and reduce it to a level dictated by discreteness in the Poisson case. We generalize the method in order to apply it to analysis of experiments searching for neutrino oscillations. We show that this technique both gives correct coverage and is powerful, while other classical techniques that have been used by neutrino oscillation search experiments fail one or both of these criteria.

physics.data-an↗