SearcharxivSearch

arXiv subjects

Ronald Meester

Publications and source records attributed to Ronald Meester.

At least 19 recordsLinked to original sources

A Bayesian framework for analyzing alleged cheating in sports through hidden codes, with applications to bridge and baseball

We develop a statistical framework to evaluate evidence of alleged cheating involving illegal signaling in sports from a forensic perspective. We explain why, instead of a frequentist procedure, a Bayesian approach is called for. We apply this framework to cases of alleged cheating in professional bridge and professional baseball. The diversity of these applications illustrates the generality of the method.

stat.AP

Reconciling common source, specific source, feature based and score based likelihood ratios

We show that the incorporation of any new piece of information allows for improved decision making in the sense that the expected costs of an optimal decision decrease (or, in boundary cases where no or not enough new information is incorporated, stays the same) whenever this is done by the appropriate update of the probabilities of the hypotheses. Versions of this result have been stated before. However, previous proofs rely on auxiliary constructions with proper scoring rules. We, instead, offer a direct and completely general proof by considering elementary properties of likelihood ratios only. We apply our results to make a contribution to the debates about the use of score based/feature based and common/specific source likelihood ratios. In the literature these are often presented as different ``LR-systems''. We argue that the difference between these is simply a matter which information is processed. There is no therefore no such thing as different ``LR-systems'', there are only differences in the processed information. In particular, despite claims to the contrary, scores can very well be used in forensic practice and we illustrate this with an extensive example in DNA kinship context.

math.ST

On the relation between likelihood ratios and p-values for testing success probabilities of Bernoulli trials

It is well known that there is no direct one-to-one relation between $p$-values and likelihood ratios or Bayes factors, since their relation crucially involves the sample size $n$. We investigate their (asymptotic) relation in a coin-tossing context where the hypotheses of interest address the success probability of the coin, and where detailed computations are possible. This leads to useful insights in the nature of $p$-values and likelihood ratios. Our results imply, for instance, that under mild conditions, a $p$-value of 0.05 cannot correspond to a likelihood ratio larger than 7.5, for any hypothesis versus a null hypothesis that the success probability has a specific value. We also show it is unlikely one can obtain a large likelihood ratio by tossing a fair coin until the number of heads deviates from the mean by several standard deviations.

math.ST

On the (ab)use of statistics in the legal case against the nurse Lucia de B

We discuss the statistics involved in the legal case of the nurse Lucia de B. in The Netherlands, 2003-2004. Lucia de B. witnessed an unusually high number of incidents during her shifts, and the question arose as to whether this could be attributed to chance. We discuss and criticise the statistical analysis of Henk Elffers, a statistician who was asked by the prosecutor to write a statistical report on the issue. We discuss several other possibilities for statistical analysis. Our main point is that several statistical models exist, leading to very different predictions, or perhaps different answers to different questions. There is no such thing as a `best' statistical analysis.

math.ST

Some reflections on the test-negative design

We discuss some philosophical, methodological and practical problems concerning the use of the test-negative design for COVID-19 vaccines. These problems limit the use of this design considerably.

stat.AP

The asymptotics of group Russian roulette

We study the group Russian roulette problem, also known as the shooting problem, defined as follows. We have $n$ armed people in a room. At each chime of a clock, everyone shoots a random other person. The persons shot fall dead and the survivors shoot again at the next chime. Eventually, either everyone is dead or there is a single survivor. We prove that the probability $p_n$ of having no survivors does not converge as $n\to\infty$, and becomes asymptotically periodic and continuous on the $\log n$ scale, with period 1.

math.PR

A behavioral interpretation of belief functions

Shafer's belief functions were introduced in the seventies of the previous century as a mathematical tool in order to model epistemic probability. One of the reasons that they were not picked up by mainstream probability was the lack of a behavioral interpretation. In this paper we provide such a behavioral interpretation, and re-derive Shafer's belief functions via a betting interpretation reminiscent of the classical Dutch Book Theorem for probability distributions. We relate our betting interpretation of belief functions to the existing literature.

math.PR

Phase transition and uniqueness of levelset percolation

The main purpose of this paper is to introduce and establish basic results of a natural extension of the classical Boolean percolation model (also known as the Gilbert disc model). We replace the balls of that model by a positive non-increasing attenuation function $l:(0,\infty) \to (0,\infty)$ to create the random field $Ψ(y)=\sum_{x\in η}l(|x-y|),$ where $η$ is a homogeneous Poisson process in ${\mathbb R}^d.$ The field $Ψ$ is then a random potential field with infinite range dependencies whenever the support of the function $l$ is unbounded. In particular, we study the level sets $Ψ_{\geq h}(y)$ containing the points $y\in {\mathbb R}^d$ such that $Ψ(y)\geq h.$ In the case where $l$ has unbounded support, we give, for any $d\geq 2,$ exact conditions on $l$ for $Ψ_{\geq h}(y)$ to have a percolative phase transition as a function of $h.$ We also prove that when $l$ is continuous then so is $Ψ$ almost surely. Moreover, in this case and for $d=2,$ we prove uniqueness of the infinite component of $Ψ_{\geq h}$ when such exists, and we also show that the so-called percolation function is continuous below the critical value $h_c$.

math.PR

Assessing forensic evidence by computing belief functions

We first discuss certain problems with the classical probabilistic approach for assessing forensic evidence, in particular its inability to distinguish between lack of belief and disbelief, and its inability to model complete ignorance within a given population. We then discuss Shafer belief functions, a generalization of probability distributions, which can deal with both these objections. We use a calculus of belief functions which does not use the much criticized Dempster rule of combination, but only the very natural Dempster-Shafer conditioning. We then apply this calculus to some classical forensic problems like the various island problems and the problem of parental identification. If we impose no prior knowledge apart from assuming that the culprit or parent belongs to a given population (something which is possible in our setting), then our answers differ from the classical ones when uniform or other priors are imposed. We can actually retrieve the classical answers by imposing the relevant priors, so our setup can and should be interpreted as a generalization of the classical methodology, allowing more flexibility. We show how our calculus can be used to develop an analogue of Bayes' rule, with belief functions instead of classical probabilities. We also discuss consequences of our theory for legal practice.

math.PR

Quantifying knowledge with a new calculus for belief functions - a generalization of probability theory

We first show that there are practical situations in for instance forensic and gambling settings, in which applying classical probability theory, that is, based on the axioms of Kolmogorov, is problematic. We then introduce and discuss Shafer belief functions. Technically, Shafer belief functions generalize probability distributions. Philosophically, they pertain to individual or shared knowledge of facts, rather than to facts themselves, and therefore can be interpreted as generalizing epistemic probability, that is, probability theory interpreted epistemologically. Belief functions are more flexible and better suited to deal with certain types of uncertainty than classical probability distributions. We develop a new calculus for belief functions which does not use the much criticized Dempster's rule of combination, by generalizing the classical notions of conditioning and independence in a natural and uncontroversial way. Using this calculus, we explain our rejection of Dempster's rule in detail. We apply the new theory to a number of examples, including a gambling example and an example in a forensic setting. We prove a law of large numbers for belief functions and offer a betting interpretation similar to the Dutch Book Theorem for probability distributions.

math.PR

Generating stationary random graphs on $\mathbb{Z}$ with prescribed i.i.d.\ degrees

Let $F$ be a probability distribution with support on the non-negative integers. Two algorithms are described for generating a stationary random graph, with vertex set $\mathbb{Z}$, so that the degrees of the vertices are i.i.d.\ random variables with distribution $F$. Focus is on an algorithm where, initially, a random number of "stubs" with distribution $F$ is attached to each vertex. Each stub is then randomly assigned a direction, left or right, and the edge configuration is obtained by pairing stubs pointing to each other, first exhausting all possible connections between nearest neighbors, then linking second nearest neighbors, and so on. Under the assumption that $F$ has finite mean, it is shown that this algorithm leads to a well-defined configuration, but that the expected length of the shortest edge of a vertex is infinite. It is also shown that any stationary algorithm for pairing stubs with random, independent directions gives infinite mean for the total length of the edges of a given vertex. Connections to the problem of constructing finitary isomorphisms between Bernoulli shifts are discussed.

math.PR

Stochastic SIR epidemics in a population with households and schools

We study the spread of stochastic SIR (Susceptible $\to$ Infectious $\to$ Recovered) epidemics in two types of structured populations, both consisting of schools and households. In each of the types, every individual is part of one school and one household. In the independent partition model, the partitions of the population into schools and households are independent of each other. This model corresponds to the well-studied household-workplace model. In the hierarchical model which we introduce here, members of the same household are also members of the same school. We introduce computable branching process approximations for both types of populations and use these to compare the probabilities of a large outbreak. The branching process approximation in the hierarchical model is novel and of independent interest. We prove by a coupling argument that if all households and schools have the same size, an epidemic spreads easier (in the sense that the number of individuals infected is stochastically larger) in the independent partition model. We also show by example that this result does not necessarily hold if households and/or schools do not all have the same size.

physics.soc-ph

Uniquely determined uniform probability on the natural numbers

In this paper, we address the problem of constructing a uniform probability measure on $\mathbb{N}$. Of course, this is not possible within the bounds of the Kolmogorov axioms and we have to violate at least one axiom. We define a probability measure as a finitely additive measure assigning probability $1$ to the whole space, on a domain which is closed under complements and finite disjoint unions. We introduce and motivate a notion of uniformity which we call weak thinnability, which is strictly stronger than extension of natural density. We construct a weakly thinnable probability measure and we show that on its domain, which contains sets without natural density, probability is uniquely determined by weak thinnability. In this sense, we can assign uniform probabilities in a canonical way. We generalize this result to uniform probability measures on other metric spaces, including $\mathbb{R}^n$.

math.PR

On central limit theorems in the random connection model

Consider a sequence of Poisson random connection models (X_n,lambda_n,g_n) on R^d, where lambda_n / n^d \to lambda > 0 and g_n(x) = g(nx) for some non-increasing, integrable connection function g. Let I_n(g) be the number of isolated vertices of (X_n,lambda_n,g_n) in some bounded Borel set K, where K has non-empty interior and boundary of Lebesgue measure zero. Roy and Sarkar [Phys. A 318 (2003), no. 1-2, 230-242] claim that (I_n(g) - E I_n(g)) / \sqrt Var I_n(g) converges in distribution to a standard normal random variable. However, their proof has errors. We correct their proof and extend the result to larger components when the connection function g has bounded support.

math.PR

The signed loop approach to the Ising model: foundations and critical point

The signed loop method is a beautiful way to rigorously study the two-dimensional Ising model with no external field. In this paper, we explore the foundations of the method, including details that have so far been neglected or overlooked in the literature. We demonstrate how the method can be applied to the Ising model on the square lattice to derive explicit formal expressions for the free energy density and two-point functions in terms of sums over loops, valid all the way up to the self-dual point. As a corollary, it follows that the self-dual point is critical both for the behaviour of the free energy density, and for the decay of the two-point functions.

math.PR

Critical densities in sandpile models with quenched or annealed disorder

We discuss various critical densities in sandpile models. The stationary density is the average expected height in the stationary state of a finite-volume model; the transition density is the critical point in the infinite-volume counterpart. These two critical densities were generally assumed to be equal, but this has turned out to be wrong for deterministic sandpile models. We show they are not equal in a quenched version of the Manna sandpile model either. In the literature, when the transition density is simulated, it is implicitly or explicitly assumed to be equal to either the so-called threshold density or the so-called critical activity density. We properly define these auxiliary densities, and prove that in certain cases, the threshold density is equal to the transition density. We extend the definition of the critical activity density to infinite volume, and prove that in the standard infinite volume sandpile, it is equal to 1. Our results should bring some order in the precise relations between the various densities.

math-ph

Forensic Identification: Database likelihood ratios and familial DNA searching

Familial Searching is the process of searching in a DNA database for relatives of a certain individual. It is well known that in order to evaluate the genetic evidence in favour of a certain given form of relatedness between two individuals, one needs to calculate the appropriate likelihood ratio, which is in this context called a Kinship Index. Suppose that the database contains, for a given type of relative, at most one related individual. Given prior probabilities for being the relative for all persons in the database, we derive the likelihood ratio for each database member in favour of being that relative. This likelihood ratio takes all the Kinship Indices between the target individual and the members of the database into account. We also compute the corresponding posterior probabilities. We then discuss two methods to select a subset from the database that contains the relative with a known probability, or at least a useful lower bound thereof. One method needs prior probabilities and yields posterior probabilities, the other does not. We discuss the relation between the approaches, and illustrate the methods with familial searching carried out in the Dutch National DNA Database.

stat.AP

Fat fractal percolation and k-fractal percolation

We consider two variations on the Mandelbrot fractal percolation model. In the k-fractal percolation model, the d-dimensional unit cube is divided in N^d equal subcubes, k of which are retained while the others are discarded. The procedure is then iterated inside the retained cubes at all smaller scales. We show that the (properly rescaled) percolation critical value of this model converges to the critical value of ordinary site percolation on a particular d-dimensional lattice as N tends to infinity. This is analogous to the result of Falconer and Grimmett that the critical value for Mandelbrot fractal percolation converges to the critical value of site percolation on the same d-dimensional lattice. In the fat fractal percolation model, subcubes are retained with probability p_n at step n of the construction, where (p_n) is a non-decreasing sequence with \prod p_n > 0. The Lebesgue measure of the limit set is positive a.s. given non-extinction. We prove that either the set of connected components larger than one point has Lebesgue measure zero a.s. or its complement in the limit set has Lebesgue measure zero a.s.

math.PR