SearcharxivSearch

arXiv subjects

Harrison B. Prosper

Publications and source records attributed to Harrison B. Prosper.

At least 19 recordsLinked to original sources

Modeling Sampling Distributions of Test Statistics with Autograd

Simulation-based inference methods that feature correct conditional coverage of confidence sets based on observations that have been compressed to a scalar test statistic require accurate modeling of either the p-value function or the cumulative distribution function (cdf) of the test statistic. If the model of the cdf, which is typically a deep neural network, is a function of the test statistic then the derivative of the neural network with respect to the test statistic furnishes an approximation of the sampling distribution of the test statistic. We explore whether this approach to modeling conditional 1-dimensional sampling distributions is a viable alternative to the probability density-ratio method, also known as the likelihood-ratio trick. Relatively simple, yet effective, neural network models are used whose predictive uncertainty is quantified through a variety of methods.

stat.ML

Amortized Simulation-Based Frequentist Inference for Tractable and Intractable Likelihoods

High-fidelity simulators that connect theoretical models with observations are indispensable tools in many sciences. When coupled with machine learning, a simulator makes it possible to infer the parameters of a theoretical model directly from real and simulated observations without explicit use of the likelihood function. This is of particular interest when the latter is intractable. In this work, we introduce a simple extension of the recently proposed likelihood-free frequentist inference (LF2I) approach that has some computational advantages. Like LF2I, this extension yields provably valid confidence sets in parameter inference problems in which a high-fidelity simulator is available. The utility of our algorithm is illustrated by applying it to three pedagogically interesting examples: the first is from cosmology, the second from high-energy physics and astronomy, both with tractable likelihoods, while the third, with an intractable likelihood, is from epidemiology.

stat.ME

Democratizing LHC Data Analysis with ADL/CutLang

Data analysis at the LHC has a very steep learning curve, which erects a formidable barrier between data and anyone who wishes to analyze data, either to study an idea or to simply understand how data analysis is performed. To make analysis more accessible, we designed the so-called Analysis Description Language (ADL), a domain specific language capable of describing the contents of an LHC analysis in a standard and unambiguous way, independent of any computing frameworks. ADL has an English-like highly human-readable syntax and directly employs concepts relevant to HEP. Therefore it eliminates the need to learn complex analysis frameworks written based on general purpose languages such as C++ or Python, and shifts the focus directly to physics. Analyses written in ADL can be run on data using a runtime interpreter called CutLang, without the necessity of programming. ADL and CutLang are designed for use by anyone with an interest in, and/or knowledge of LHC physics, ranging from experimentalists and phenomenologists to non-professional enthusiasts. ADL/CutLang are originally designed for research, but are also equally intended for education and public use. This approach has already been employed to train undergraduate students with no programming experience in LHC analysis in two dedicated schools in Turkey and Vietnam, and is being adapted for use with LHC Open Data. Moreover, work is in progress towards piloting an educational module in particle physics data analysis for high school students and teachers. Here, we introduce ADL and CutLang and present the educational activities based on these practical tools.

physics.ed-ph

Analysis Description Language: A DSL for HEP Analysis

We propose to adopt a declarative domain specific language for describing the physics algorithm of a high energy physics (HEP) analysis in a standard and unambiguous way decoupled from analysis software frameworks, and argue that this approach provides an accessible and sustainable environment for analysis design, use and preservation. Prototype of such a language called Analysis Description Language (ADL) and its associated tools are being developed and applied in various HEP physics studies. We present the motivations for using a DSL, design principles of ADL and its runtime interpreter CutLang, along with current physics studies based on this approach. We also outline ideas and prospects for the future. Recent physics studies, hands-on workshops and surveys indicate that ADL is a feasible and effective approach with many advantages and benefits, and offers a direction to which the HEP field should give serious consideration.

hep-ph

Implicit Quantile Neural Networks for Jet Simulation and Correction

Reliable modeling of conditional densities is important for quantitative scientific fields such as particle physics. In domains outside physics, implicit quantile neural networks (IQN) have been shown to provide accurate models of conditional densities. We present a successful application of IQNs to jet simulation and correction using the tools and simulated data from the Compact Muon Solenoid (CMS) Open Data portal.

physics.comp-ph

Publishing statistical models: Getting the most out of particle physics experiments

The statistical models used to derive the results of experimental analyses are of incredible scientific value and are essential information for analysis preservation and reuse. In this paper, we make the scientific case for systematically publishing the full statistical models and discuss the technical developments that make this practical. By means of a variety of physics cases -- including parton distribution functions, Higgs boson measurements, effective field theory interpretations, direct searches for new physics, heavy flavor physics, direct dark matter detection, world averages, and beyond the Standard Model global fits -- we illustrate how detailed information on the statistical modelling can enhance the short- and long-term impact of experimental results.

hep-ph

Recent advances in ADL, CutLang and adl2tnm

This paper presents an overview and features of an Analysis Description Language (ADL) designed for HEP data analysis. ADL is a domain specific, declarative language that describes the physics content of an analysis in a standard and unambiguous way, independent of any computing frameworks. It also describes infrastructures that render ADL executable, namely CutLang, a direct runtime interpreter (originally also a language), and adl2tnm, a transpiler converting ADL into C++ code. In ADL, analyses are described in human readable plain text files, clearly separating object, variable and event selection definitions in blocks, with a syntax that includes mathematical and logical operations, comparison and optimisation operators, reducers, four-vector algebra and commonly used functions. Recent studies demonstrate that adapting the ADL approach has numerous benefits for the experimental and phenomenological HEP communities. These include facilitating the abstraction, design, optimization, visualization, validation, combination, reproduction, interpretation and overall communication of the analysis contents and long term preservation of the analyses beyond the lifetimes of experiments. Here we also discuss some of the current ADL applications in physics studies and future prospects based on static analysis and differentiable programming.

physics.comp-ph

Analysis Description Languages for the LHC

An analysis description language is a domain specific language capable of describing the contents of an LHC analysis in a standard and unambiguous way, independent of any computing framework. It is designed for use by anyone with an interest in, and knowledge of, LHC physics, i.e., experimentalists, phenomenologists and other enthusiasts. Adopting analysis description languages would bring numerous benefits for the LHC experimental and phenomenological communities ranging from analysis preservation beyond the lifetimes of experiments or analysis software to facilitating the abstraction, design, visualization, validation, combination, reproduction, interpretation and overall communication of the analysis contents. Here, we introduce the analysis description language concept and summarize the current efforts ongoing to develop such languages and tools to use them in LHC analyses.

hep-ph

Optimizing Event Selection with the Random Grid Search

The random grid search (RGS) is a simple, but efficient, stochastic algorithm to find optimal cuts that was developed in the context of the search for the top quark at Fermilab in the mid-1990s. The algorithm, and associated code, have been enhanced recently with the introduction of two new cut types, one of which has been successfully used in searches for supersymmetry at the Large Hadron Collider. The RGS optimization algorithm is described along with the recent developments, which are illustrated with two examples from particle physics. One explores the optimization of the selection of vector boson fusion events in the four-lepton decay mode of the Higgs boson and the other optimizes SUSY searches using boosted objects and the razor variables.

hep-ph

Practical Statistics for Particle Physicists

These lectures introduce the basic ideas and practices of statistical analysis for particle physicists, using a real-world example to illustrate how the abstractions on which statistics is based are translated into practical application.

physics.data-an

Prospect for measuring the CP phase in the $hττ$ coupling at the LHC

The search for a new source of CP violation is one of the most important endeavors in particle physics. A particularly interesting way to perform this search is to probe the CP phase in the $hττ$ coupling, as the phase is currently completely unconstrained by all existing data. Recently, a novel variable $Θ$ was proposed for measuring the CP phase in the $hττ$ coupling through the $τ^\pm \to π^\pm π^0 ν$ decay mode. We examine two crucial questions that the real LHC detectors must face, namely, the issue of neutrino reconstruction and the effects of finite detector resolution. For the former, we find strong evidence that the collinear approximation is the best for the $Θ$ variable. For the latter, we find that the angular resolution is actually not an issue even though the reconstruction of $Θ$ requires resolving the highly collimated $π^\pm$'s and $π^0$'s from the $τ$ decays. Instead, we find that it is the missing transverse energy resolution that significantly limits the LHC reach for measuring the CP phase via $Θ$. With the current missing energy resolution, we find that with $\sim 1000\,\textrm{fb}^{-1}$ the CP phase hypotheses $Δ= 0^\circ$ (the standard model value) and $Δ= 90^\circ$ can be distinguished, at most, at the 95\% confidence level.

hep-ph

Discovery potential for heavy t-tbar resonances in dilepton+jets final states

We examine the prospects for probing heavy top quark-antiquark (t-tbar) resonances at the upgraded LHC in pp collisions at $\root_s = 14 TeV. Heavy t-tbar resonances (Z' bosons) are predicted by several theories that go beyond the standard model. We consider scenarios in which each top quark decays leptonically, either to an electron or a muon, and the data sets correspond to integrated luminosities of \int L dt = 300 /fb and \int L dt = 3000 /fb. We present the expected 5-sigma discovery potential for a Z' resonance as well as the expected upper limits at 95% C.L. on the Z' production cross section and mass in the absence of a discovery.

hep-ex

Priors for New Physics

The interpretation of data in terms of multi-parameter models of new physics, using the Bayesian approach, requires the construction of multi-parameter priors. We propose a construction that uses elements of Bayesian reference analysis. Our idea is to initiate the chain of inference with the reference prior for a likelihood function that depends on a single parameter of interest that is a function of the parameters of the physics model. The reference posterior density of the parameter of interest induces on the parameter space of the physics model a class of posterior densities. We propose to continue the chain of inference with a particular density from this class, namely, the one for which indistinguishable models are equiprobable and use it as the prior for subsequent analysis. We illustrate our method by applying it to the constrained minimal supersymmetric Standard Model and two non-universal variants of it.

physics.data-an

Varying-G Cosmology with Type Ia Supernovae

The observation that Type Ia supernovae are fainter than expected given their red shifts has led to the conclusion that the expansion of the universe is accelerating. The widely accepted hypothesis is that this acceleration is caused by a cosmological constant or, more generally, some dark energy field that pervades the universe. This hypothesis presents a challenge to physics so severe that one is motivated to explore alternative explanations. In this paper, we explore whether the data from Type Ia supernovae can be explained with an idea that is almost as old as that of the cosmological constant, namely, that the strength of gravity varies on a cosmic timescale. This topic is an ideal one for investigation by an undergraduate physics major because the entire chain of reasoning from models to data analysis is well within the mathematical and conceptual sophistication of a motivated undergraduate.

astro-ph.CO

Reference priors for high energy physics

Bayesian inferences in high energy physics often use uniform prior distributions for parameters about which little or no information is available before data are collected. The resulting posterior distributions are therefore sensitive to the choice of parametrization for the problem and may even be improper if this choice is not carefully considered. Here we describe an extensively tested methodology, known as reference analysis, which allows one to construct parametrization-invariant priors that embody the notion of minimal informativeness in a mathematically well-defined sense. We apply this methodology to general cross section measurements and show that it yields sensible results. A recent measurement of the single top quark cross section illustrates the relevant techniques in a realistic situation.

stat.AP

Probability and Statistical Inference

These lectures introduce key concepts in probability and statistical inference at a level suitable for graduate students in particle physics. Our goal is to paint as vivid a picture as possible of the concepts covered.

physics.data-an

Strategy for discovering a low-mass Higgs boson at the Fermilab Tevatron

We have studied the potential of the CDF and DZero experiments to discover a low-mass Standard Model Higgs boson, during Run II, via the processes $p\bar{p}$ -> WH -> $\ellνb\bar{b}$, $p\bar{p}$ -> ZH -> $\ell^{+}\ell^{-}b\bar{b}$ and $p\bar{p}$ -> ZH ->$ν\barν b\bar{b}$. We show that a multivariate analysis using neural networks, that exploits all the information contained within a set of event variables, leads to a significant reduction, with respect to {\em any} equivalent conventional analysis, in the integrated luminosity required to find a Standard Model Higgs boson in the mass range 90 GeV/c**2 < M_H < 130 GeV/c**2. The luminosity reduction is sufficient to bring the discovery of the Higgs boson within reach of the Tevatron experiments, given the anticipated integrated luminosities of Run II, whose scope has recently been expanded.

hep-ph