SearcharxivSearch

arXiv subjects

Lutz Mattner

Publications and source records attributed to Lutz Mattner.

At least 19 recordsLinked to original sources

Teachable normal approximations to binomial and related probabilities or confidence bounds

For the usual normal approximations to binomial, hypergeometric, or Poisson interval probabilities, we collect some simple but then reasonably sharp error bounds. For the Clopper-Pearson~(1934) binomial confidence bounds, we present, following Michael Short's~(2023) approach, bounds similar to, but necessarily more complicated than, Lagrange's (1776) success rate plus/minus normal quantile times estimated standard deviation. The bounds, as presented here in four theorems, should be teachable, to people ranging from sufficiently advanced high school pupils to university students in mathematics or statistics: For understanding most of the proposed approximation results, it should suffice to know binomial laws, their means and variances, and the standard normal distribution function, but not necessarily the concept of a corresponding normal random variable. Accompanying technical remarks, references, and proofs are meant for assuring teachers or for stimulating further research. Of the proposed approximations, some are essentially well-known at least to experts, and some are based on teaching experience and research at Trier University.

stat.OT

A sharper Lyapunov-Katz central limit error bound for i.i.d. summands Zolotarev-close to normal

We prove a central limit error bound for convolution powers of laws with finite moments of order $r \in \mathopen]2,3\mathclose]$, taking a closeness of the laws to normality into account. Up to a universal constant, this generalises the case of $r=3$ of the sharpening of the Berry (1941) - Esseen (1942) theorem obtained by Mattner (2024), namely by sharpening here the Katz (1963) error bound for the i.i.d. case of Lyapunov's (1901) theorem. Our proof uses a partial generalisation of the theorem of Senatov and Zolotarev (1986) used for the earlier special case. A result more general than our main one could be obtained by using instead another theorem of Senatov (1980), but unfortunately an auxiliary inequality used in the latter's proof is wrong.

math.PR

A convolution inequality, yielding a sharper Berry-Esseen theorem for summands Zolotarev-close to normal

The classical Berry-Esseen error bound, for the normal approximation to the law of a sum of independent and identically distributed random variables, is here improved by replacing the standardised third absolute moment by a weak norm distance to normality. We thus sharpen and simplify two results of Ulyanov (1976) and of Senatov (1998), each of them previously optimal, in the line of research initiated by Zolotarev (1965) and Paulauskas (1969). Our proof is based on a seemingly incomparable normal approximation theorem of Zolotarev (1986), combined with our main technical result: The Kolmogorov distance (supremum norm of difference of distribution functions) between a convolution of two laws and a convolution of two Lipschitz laws is bounded homogeneously of degree 1 in the pair of the Wasserstein distances (L$^1$ norms of differences of distribution functions) of the corresponding factors, and also, inessentially for the present application, in the pair of the Lipschitz constants. Side results include a short introduction to $ζ$ norms on the real line, simpler inequalities for various probability distances, slight improvements of the theorem of Zolotarev (1986) and of a lower bound theorem of Bobkov, Chistyakov and Götze (2012), an application to sampling from finite populations, auxiliary results on rounding and on winsorisation, and computations of a few examples. The introductory section in particular is aimed at analysts in general rather than specialists in probability approximations.

math.PR

Extreme expectations of Bernoulli convolutions given their first few moments are attained at shifted convolutions of as few binomials

A result of Chebyshev (1864) and Hoeffding1956}, on bounding an expectation of a given function with respect to a Bernoulli convolution (also called Poisson binomial law, or law of the number of successes in independent trials) with any given first moment, is here generalised to the case of any given first few moments, as indicated in the title. A nonprobabilistic, and perhaps more obvious, reformulation is: Every permutation invariant and separately affine-linear function of $n$ real variables $x_i\in[a,b]$ assumes its extremal values given the power sums $\sum_{i=1}^nx_i^1,\ldots, \sum_{i=1}^nx_i^r$ at vectors $x$ with at most $r$ coordinate values different from $a$ and $b$.

math.PR

An optimal Berry-Esseen type theorem for integrals of smooth functions

We prove a Berry-Esseen type inequality for approximating expectations of sufficiently smooth functions $f$, like $f=|\cdot|^3$, with respect to standardized convolutions of laws $P_1,\ldots, P_n$ on the real line by corresponding expectations based on symmetric two-point laws $Q_1,\ldots,Q_n$ isoscedastic to the $P_i$. Equality is attained for every possible constellation of the Lipschitz constant $\|f"\|^{}_{\mathrm{L}}$ and the variances and the third centred absolute moments of the $P_i$. The error bound is strictly smaller than $\frac 16$ times the Lyapunov ratio times $\|f"\|^{}_{\mathrm{L}}$, and tends to zero also if $n$ is fixed and the third standardized absolute moments of the $P_i$ tend to one. In the homoscedastic case of equal variances of the $P_i$, and hence in particular in the i.i.d. case, the approximating law is a standardized symmetric binomial one. The inequality is strong enough to yield for some constellations, in particular in the i.i.d. case with $n$ large enough given the standardized third absolute moment of $P_1$, an improvement of a more classical and already optimal Berry-Esseen type inequality of Tyurin (2009). Auxiliary results presented include some inequalities either purely analytical or concerning Zolotarev's $ζ$-metrics, and some binomial moment calculations.

math.PR

The medians for exponential families and the normal law

Let $P$ a probability on the real line generating a natural exponential family $(P_t)_{t\in \R}$. We show that $t$ is a median of $P_t$ for all $t$ only if $P$ is the standard Gaussian law $N(0.1).$ The proof is based on the Choquet Deny equation.

math.PR

Confidence intervals for average success probabilities

We provide Buehler-optimal one-sided and some valid two-sided confidence intervals for the average success probability of a possibly inhomogeneous fixed length Bernoulli chain, based on the number of observed successes. Contrary to some claims in the literature, the one-sided Clopper-Pearson intervals for the homogeneous case are not completely robust here, not even if applied to hypergeometric estimation problems.

math.ST

Partially complete sufficient statistics are jointly complete

The theory of the basic statistical concept of (Lehmann-Scheffé-)completeness is perfected by providing the theorem indicated in the title and previously overlooked for several decades. Relations to earlier results are discussed and illustrating examples are presented. Of the two proofs offered for the main result, the first is direct and short, following the prototypical example of Landers and Rogge (1976), and the second is very short and purely statistical, utilizing the basic theory of optimal unbiased estimation in the little known version completed by Schmetterer and Strasser (1974).

math.ST

On normal approximations to symmetric hypergeometric laws

The Kolmogorov distances between a symmetric hypergeometric law with standard deviation $σ$ and its usual normal approximations are computed and shown to be less than $1/(\sqrt{8π}\,σ)$, with the order $1/σ$ and the constant $1/\sqrt{8π}$ being optimal. The results of Hipp and Mattner (2007) for symmetric binomial laws are obtained as special cases. Connections to Berry-Esseen type results in more general situations concerning sums of simple random samples or Bernoulli convolutions are explained. Auxiliary results of independent interest include rather sharp normal distribution function inequalities, a simple identifiability result for hypergeometric laws, and some remarks related to Lévy's concentration-variance inequality.

math.PR

Confidence bounds for the sensitivity lack of a less specific diagnostic test, without gold standard

We consider the problem of comparing two diagnostic tests based on a sample of paired test results without true state determinations, in cases where the second test can reasonably be assumed to be at least as specific as the first. For such cases, we provide two informative confidence bounds: A lower one for the prevalence times the sensitivity gain of the second test with respect to the first, and an upper one for the sensitivity of the first test. Neither conditional independence of the two tests nor perfectness of any of them needs to be assumd. An application of the proposed confidence bounds to a sample of 256 pairs of laboratory test results for toxigenic Clostridium difficile provides evidence for a dramatic sensitivity gain through first appropriately culturing Clostridium difficile from stool samples before applying an enzyme-immuno-assay.

stat.AP

Combining individually valid and conditionally i.i.d. P-variables

For a given testing problem, let $U_1,...,U_n$ be individually valid and conditionally on the data i.i.d.\ P-variables (often called P-values). For example, the data could come in groups, and each $U_i$ could be based on subsampling just one datum from each group in order to satisfy an independence assumption under the hypothesis. The problem is then to deterministically combine the $U_i$ into a valid summary P-variable. Restricting here our attention to functions of a given order statistic $U_{k:n}$ of the $U_i$, we compute the function $f_{n,k}$ which is smallest among all increasing functions $f$ such that $f(U_{k:n})$ is always a valid P-variable under the stated assumptions. Since $f_{n,k}(u)\le 1\wedge (\frac {n}{k} u)$, with the right hand side being a good approximation for the left when $k$ is large, one may in particular always take the minimum of 1 and twice the left sample median of the given P-variables. We sketch the original application of the above in a recent study of associations between various primate species by Astaras et al.

stat.ME

Stochastic ordering of classical discrete distributions

For several pairs $(P,Q)$ of classical distributions on $\N_0$, we show that their stochastic ordering $P\leq_{st} Q$ can be characterized by their extreme tail ordering equivalent to $ P(\{k_\ast \})/Q(\{k_\ast\}) \le 1 \le \lim_{k\to k^\ast} P(\{k\})/Q(\{k\})$, with $k_\ast$ and $k^\ast$ denoting the minimum and the supremum of the support of $P+Q$, and with the limit to be read as $P(\{k^\ast\})/Q(\{k^\ast\})$ for $k^\ast$ finite. This includes in particular all pairs where $P$ and $Q$ are both binomial ($b_{n_1,p_1} \leq_{st} b_{n_2,p_2}$ if and only if $n_1\le n_2$ and $(1-p_1)^{n_1}\ge(1-p_2)^{n_2}$, or $p_1=0$), both negative binomial ($b^-_{r_1,p_1}\leq_{st} b^-_{r_2,p_2}$ if and only if $p_1\geq p_2$ and $p_1^{r_1}\geq p_2^{r_2}$), or both hypergeometric with the same sample size parameter. The binomial case is contained in a known result about Bernoulli convolutions, the other two cases appear to be new. The emphasis of this paper is on providing a variety of different methods of proofs: (i) half monotone likelihood ratios, (ii) explicit coupling, (iii) Markov chain comparison, (iv) analytic calculation, and (v) comparison of Levy measures. We give four proofs in the binomial case (methods (i)-(iv)) and three in the negative binomial case (methods (i), (iv) and (v)). The statement for hypergeometric distributions is proved via method (i).

math.PR

One optional observation inflates $α$ by $100/\sqrt{n}$ per cent

For one-sample level $α$ tests $ψ_m$ based on independent observations $X_1,...,X_m$, we prove an asymptotic formula for the actual level of the test rejecting if at least one of the tests $ψ_{n},...,ψ_{n+k}$ would reject. For $k=1$ and usual tests at usual levels $α$, the result is approximately summarized by the title of this paper. Our method of proof, relying on some second order asymptotic statistics as developed by Pfanzagl and Wefelmeyer, might also be useful for proper sequential analysis. A simple and elementary alternative proof is given for $k=1$ in the special case of the Gauss test.

math.ST

Optimal L$^1$-bounds for submartingales

The optimal function $f$ satisfying $$ \mathbb{E} |\sum_{1}^n X_i | \ge f(\mathrbb{E}|X_1|,...,\mathbb{E}|X_n|) $$ for every martingale $(X_1,X_1+X_2, ...,\sum_{i=1}^n X_i)$ is shown to be given by $$ f(a) = \max \Big\{a_k-\sum_{i=1}^{k-1} a_i\Big\}_{k=1}^n \cup \Big\{\frac {a_k}2\Big\}_{k=3}^n $$ for $a\in{[0,\infty[}^n_{}$. A similar result is obtained for submartingales $(0,X_1,X_1+X_2,..., \sum_{i=1}^n X_i)$. The optimality proofs use a convex-analytic comparison lemma of independent interest.

math.PR

A shorter proof of Kanter's Bessel function concentration bound

We give a shorter proof of Kanter's (1976) sharp Bessel function bound for concentrations of sums of independent symmetric random vectors. We provide sharp upper bounds for the sum of modified Bessel functions $I_0(x)+I_1(x)$, which might be of independent interest. Corollaries improve concentration or smoothness bounds for sums of independent random variables due to Cekanavicius & Roos (2006), Roos (2005), Barbour & Xia 1999), and Le Cam (1986).

math.PR