Searcharxiv⌕ Search

arXiv subjects

Giulio D'Agostini

Publications and source records attributed to Giulio D'Agostini.

13 recordsLinked to original sources

What is the probability that a vaccinated person is shielded from Covid-19? A Bayesian MCMC based reanalysis of published data with emphasis on what should be reported as 'efficacy'

Based on the information communicated in press releases, and finally published towards the end of 2020 by Pfizer, Moderna and AstraZeneca, we have built up a simple Bayesian model, in which the main quantity of interest plays the role of {\em vaccine efficacy} (`$ε$'). The resulting Bayesian Network is processed by a Markov Chain Monte Carlo (MCMC), implemented in JAGS interfaced to R via rjags. As outcome, we get several probability density functions (pdf's) of $ε$, each conditioned on the data provided by the three pharma companies. The result is rather stable against large variations of the number of people participating in the trials and it is `somehow' in good agreement with the results provided by the companies, in the sense that their values correspond to the most probable value (`mode') of the pdf's resulting from MCMC, thus reassuring us about the validity of our simple model. However we maintain that the number to be reported as `vaccine efficacy' should be the mean of the distribution, rather than the mode, as it was already very clear to Laplace about 250 years ago (its `rule of succession' follows from the simplest problem of the kind). This is particularly important in the case in which the number of successes equals the numbers of trials, as it happens with the efficacy against `severe forms' of infection, claimed by Moderna to be 100%. The implication of the various uncertainties on the predicted number of vaccinated infectees is also shown, using both MCMC and approximated formulae.

stat.AP↗

Ratio of counts vs ratio of rates in Poisson processes

The often debated issue of `ratios of small numbers of events' is approached from a probabilistic perspective, making a clear distinction between the predictive problem (forecasting numbers of events we might count under well stated assumptions, and therefore of their ratios) and inferential problem (learning about the relevant parameters of the related probability distribution, in the light of the observed number of events). The quantities of interests and their relations are visualized in a graphical model (`Bayesian network'), very useful to understand how to approach the problem following the rules of probability theory. In this paper, written with didactic intent, we discuss in detail the basic ideas, however giving some hints of how real life complications, like (uncertain) efficiencies and possible background and systematics, can be included in the analysis, as well as the possibility that the ratio of rates might depend on some physical quantity. The simple models considered in this paper allow to obtain, under reasonable assumptions, closed expressions for the rates and their ratios. Monte Carlo methods are also used, both to cross check the exact results and to evaluate by sampling the ratios of counts in the cases in which large number approximation does not hold. In particular it is shown how to make approximate inferences using a Markov Chain Monte Carlo using JAGS/rjags. Some examples of R and JAGS code are provided.

stat.ME↗

The Gauss' Bayes Factor

In 'Theoria motus corporum coelestium in sectionibus conicis solem ambientum' Gauss presents, as a theorem and with emphasis, the rule to update the ratio of probabilities of complementary hypotheses, in the light of an observed event which could be due to either of them. Although he focused on a priori equally probable hypotheses, in order to solve the problem on which he was interested in, the theorem can be easily extended to the general case. But, curiously, I have not been able to find references to his result in the literature.

math.HO↗

Checking individuals and sampling populations with imperfect tests

In the last months, due to the emergency of Covid-19, questions related to the fact of belonging or not to a particular class of individuals (`infected or not infected'), after being tagged as `positive' or `negative' by a test, have never been so popular. Similarly, there has been strong interest in estimating the proportion of a population expected to hold a given characteristics (`having or having had the virus'). Taking the cue from the many related discussions on the media, in addition to those to which we took part, we analyze these questions from a probabilistic perspective (`Bayesian'), considering several effects that play a role in evaluating the probabilities of interest. The resulting paper, written with didactic intent, is rather general and not strictly related to pandemics: the basic ideas of Bayesian inference are introduced and the uncertainties on the performances of the tests are treated using the metrological concepts of `systematics', and are propagated into the quantities of interest following the rules of probability theory; the separation of `statistical' and `systematic' contributions to the uncertainty on the inferred proportion of infectees allows to optimize the sample size; the role of `priors', often overlooked, is stressed, however recommending the use of `flat priors', since the resulting posterior distribution can be `reshaped' by an `informative prior' in a later step; details on the calculations are given, also deriving useful approximated formulae, the tough work being however done with the help of direct Monte Carlo simulations and Markov Chain Monte Carlo, implemented in R and JAGS (relevant code provided in appendix).

q-bio.PE↗

On a curious bias arising when the $\sqrt{χ^2/ν}$ scaling prescription is first applied to a sub-sample of the individual results

As it is well known, the standard deviation of a weighted average depends only on the individual standard deviations, but not on the dispersion of the values around the mean. This property leads sometimes to the embarrassing situation in which the combined result 'looks' somehow at odds with the individual ones. A practical way to cure the problem is to enlarge the resulting standard deviation by the $\sqrt{χ^2/ν}$ scaling, a prescription employed with arbitrary criteria on when to apply it and which individual results to use in the combination. But the `apparent' discrepancy between the combined result and the individual ones often remains. Moreover this rule does not affect the resulting `best value', even if the pattern of the individual results is highly skewed. In addition to these reasons of dissatisfaction, shared by many practitioners, the method causes another issue, recently noted on the published measurements of the charged kaon mass. It happens in fact that, if the prescription is applied twice, i.e. first to a sub-sample of the individual results and subsequently to the entire sample, then a bias on the result of the overall combination is introduced. The reason is that the prescription does not guaranty statistical sufficiency, whose importance is reminded in this script, written with a didactic spirit, with some historical notes and with a language to which most physicists are accustomed. The conclusion contains general remarks on the effective presentation of the experimental findings and a pertinent puzzle is proposed in the Appendix.

physics.data-an↗

Skeptical combination of experimental results using JAGS/rjags with application to the K$^{\pm}$ mass determination

The question of how to combine experimental results that `appear' to be in mutual disagreement, treated in detail years ago in a previous paper, is revisited. The first novelty of the present note is the explicit use of graphical models, in order to make the deterministic and probabilistic links between the variables of interest more evident. Then, instead of aiming for results in closed formulae, the integrals of interest are evaluated by {\em Markov Chain Monte Carlo} (MCMC) sampling, with the algorithms (typically Gibbs Sampler) implemented in the package JAGS ("Just Another Gibbs Sampler"). For convenience, the JAGS functions are called from R scripts, thus gaining the advantage given by the rich collection of mathematical, statistical and graphical functions included in the R installation. The results of the previous paper are thus easily re-obtained and the method is applied to the determination of the charged kaon mass. This note, based on lectures to PhD students and young researchers has been written with a didactic touch, and the relevant JAGS/rjags code is provided. (A curious bias arising from the sequential application of the $\sqrt{χ^2/ν}$ scaling prescription to 'apparently' discrepant results, found here, will be discussed in more detail in a separate paper.)

physics.data-an↗

Talking about Probability, Inference and Decisions. Part 1: The Witches of Bayes

In October 2017 the Italian National Institute of Statistics (ISTAT), Italy's body for official statistics, has published the book of fairy tales Le streghe di Bayes (The witches of Bayes) written by ISTAT staff members with the commendable aim of introducing statistical and probabilistic reasoning to children. In this paper the fairy tale which gives the name to the book is analyzed in a dialog between three teachers with different background and expertise. The outcomes are definitively discouraging, especially when the story is compared to the appendix of the book, in which the teaching power of every story is indeed explained (as a matter of fact, without the appendix the fairy tale of the witches seemed to be written with the purpose of make the 'Bayesians', meant as the villagers from 'Bayes', ridiculous). In fact the fairy tale of the witches does not contain any Bayesian reasoning, the suggested decision strategy is simply wrong and the story does not even seem to be easily modifiable (besides the trivial correction of the decision strategy) in order to make it usable as a teaching tool. As it happens in real dialogues, besides the fairy tale in question, the dialogue touches several issues somehow related to the story and concerning probability, inference, prediction and decision making. The present paper is an indirect response to the invitation by the ISBA bulletin to comment on the fairy tale.

math.HO↗

More lessons from the six box toy experiment

Following a paper in which the fundamental aspects of probabilistic inference were introduced by means of a toy experiment, details of the analysis of simulated long sequences of extractions are shown here. In fact, the striking performance of probability-based inference and forecasting, compared to those obtained by simple `rules', might impress those practitioners who are usually underwhelmed by the philosophical foundation of the different methods. The analysis of the sequences also shows how the smallness of the probability of what has been actually observed, given the hypotheses of interest, is irrelevant for the purpose of inference.

math.HO↗

Probability, propensity and probabilities of propensities (and of probabilities)

The process of doing Science in condition of uncertainty is illustrated with a toy experiment in which the inferential and the forecasting aspects are both present. The fundamental aspects of probabilistic reasoning, also relevant in real life applications, arise quite naturally and the resulting discussion among non-ideologized, free-minded people offers an opportunity for clarifications.

math.HO↗

The Waves and the Sigmas (To Say Nothing of the 750 GeV Mirage)

This paper shows how p-values do not only create, as well known, wrong expectations in the case of flukes, but they might also dramatically diminish the `significance' of most likely genuine signals. As real life examples, the 2015 first detections of gravitational waves are discussed. The March 2016 statement of the American Statistical Association, warning scientists about interpretation and misuse of p-values, is also reminded and commented. (The paper is complemented with some remarks on past, recent and future claims of discoveries based on sigmas from Particles Physics.)

physics.data-an↗

Learning about probabilistic inference and forecasting by playing with multivariate normal distributions

The properties of the normal distribution under linear transformation, as well the easy way to compute the covariance matrix of marginals and conditionals, offer a unique opportunity to get an insight about several aspects of uncertainties in measurements. The way to build the overall covariance matrix in a few, but conceptually relevant cases is illustrated: several observations made with (possibly) different instruments measuring the same quantity; effect of systematics (although limited to offset, in order to stick to linear models) on the determination of the 'true value', as well in the prediction of future observations; correlations which arise when different quantities are measured with the same instrument affected by an offset uncertainty; inferences and predictions based on averages; inference about constrained values; fits under some assumptions (linear models with known standard deviations). Many numerical examples are provided, exploiting the ability of the R language to handle large matrices and to produce high quality plots. Some of the results are framed in the general problem of 'propagation of evidence', crucial in analyzing graphical models of knowledge.

physics.data-an↗

Bertrand `paradox' reloaded (with details on transformations of variables, an introduction to Monte Carlo simulation and an inferential variation of the problem)

This note is mainly to point out, if needed, that uncertainty about models and their parameters has little to do with a `paradox'. The proposed `solution' is to formulate practical questions instead of seeking refuge into abstract principles. (And, in order to be concrete, some details on how to calculate the probability density functions of the chord lengths are provided, together with some comments on simulations and an appendix on the inferential aspects of the problem.)

physics.data-an↗

Bayesian model comparison applied to the Explorer-Nautilus 2001 coincidence data

Bayesian reasoning is applied to the data by the ROG Collaboration, in which gravitational wave (g.w.) signals are searched for in a coincidence experiment between Explorer and Nautilus. The use of Bayesian reasoning allows, under well defined hypotheses, even tiny pieces of evidence in favor of each model to be extracted from the data. The combination of the data of several experiments can therefore be performed in an optimal and efficient way. Some models for Galactic sources are considered and, within each model, the experimental result is summarized with the likelihood rescaled to the insensitivity limit value (``${\cal R}$ function''). The model comparison result is given in in terms of Bayes factors, which quantify how the ratio of beliefs about two alternative models are modified by the experimental observation

gr-qc↗