SearcharxivSearch

arXiv subjects

Nick W. Koning

Publications and source records attributed to Nick W. Koning.

12 recordsLinked to original sources

Axioms for testing with data-dependent levels and e-values

The emerging literature on hypothesis testing with data-dependent and post-hoc significance levels relies on a particular extension of the Type-I error to data-dependent levels. Existing arguments for this extension are heuristic, and primarily motivated by a resulting connection to the e-value. Our first contribution is to show that it is uniquely characterized by three natural axioms. Our second contribution is to reverse the connection to the e-value, to show that three analogous axioms support a recently proposed decision-theoretic definition of the e-value. If one wishes to distinguish non-rejections at different decisions, we show three simple additional axioms suffice. Finally, we show that the relationship between e-values and post-hoc testing requires just two of these axioms.

math.ST

How e-values generalize hypothesis testing: a Neyman-Pearson lemma for the e-value

While e-values are swiftly rising in prominence, their formal relationship to classical hypothesis testing remains unsettled. We develop a unified decision-theoretic framework by viewing the e-value as a multi-decision generalization of a hypothesis test. Replacing power by the expected utility of evidence, we derive a Neyman-Pearson lemma for e-values, with classical Neyman-Pearson tests and log-optimal e-values arising from particular utility functions. Our main technical contribution is to study such expected utility-optimal e-values without assumptions on the composite null hypothesis and simple alternative. As a corollary, we obtain an assumption-free form of the classical Neyman-Pearson lemma for composite null hypotheses, showing power-optimal tests exist and have a likelihood-ratio form.

math.ST

The E-measure

We introduce the E-measure: a measure-like generalization of the E-value to a class of hypotheses. Unlike classical measures, E-measures are closed under infimums instead of addition. They arise from a compatibility axiom with logical implications, that there should be at least as much evidence against more specific hypotheses. We show that E-measures are the only non-dominated such objects, if the hypothesis class is closed under intersections. We propose to use the E-measure to present all the relevant evidence for a problem, where the relevance is captured by the choice of hypothesis class. We showcase this by applying the E-measure to decision making, inducing a hypothesis class from the uncertain consequences of decisions. This results in uniform E-consequence bounds on decisions, which nest high-probability loss bounds. Correcting for multiplicity, we consider 'familywise evidence' and 'false evidence rate' control, generalizing from errors and discoveries to continuous evidence. Remarkably, E-measures control these without multiplicity correction if the hypothesis class is intersection-closed. Moreover, we obtain a 'frequentist' notion of updating from E-prior to E-posterior. Abstracting the notion of a 'hypothesis', we advocate for using E-measures for any unknown quantity, leading to predictive E-measures.

math.ST

Fuzzy Prediction Sets: Conformal Prediction with E-values

Prediction sets offer a binary inclusion/exclusion for each element at the same fixed confidence level. We generalize to fuzzy prediction sets, which exclude elements at their own data-driven confidence level. Our key insight is that a fuzzy prediction set \emph{is} an e-value, capturing precisely what e-values bring to predictive inference. Fuzzy prediction sets inherit the merging properties of their e-value, offer richer guarantees to decision-makers. We also show in what sense optimal e-values give rise to optimal (fuzzy) prediction sets. We apply our results to conformal prediction, deriving optimal fuzzy conformal prediction sets, and characterizing in what sense classical conformal prediction is optimal.

math.ST

Equivalence testing with data-dependent and post-hoc equivalence margins

Equivalence testing compares the hypothesis that an effect $μ$ is large against the alternative that it is negligible. Here, `large' is classically expressed as being larger than some `equivalence margin' $Δ$. A longstanding problem is that this margin must be specified but can rarely be objectively justified in practice. We lay the foundation for an alternative paradigm, arguing to instead report a data-dependent margin $\widehatΔ_α$ that bounds the true effect $μ$ with probability $1 - α$. Our key argument is that $\widehatΔ_α$ is more useful than a test outcome at a fixed margin $Δ$, as measured by the guarantees it offers to decision makers. We generalize this to a curve of margins $α\mapsto \widehatΔ_α$, uniformly valid under the post-hoc selection of the margin. These ideas rely on e-values, which we derive for models that are strictly totally positive of order 3, nesting the classical z-test and t-test settings.

math.ST

Measuring Evidence against Exchangeability and Group Invariance with E-values

We study e-values for quantifying evidence against exchangeability and general invariance of a random variable under a compact group. We start by characterizing such e-values, and explaining how they nest traditional group invariance tests as a special case. We show they can be easily designed for an arbitrary test statistic, and computed through Monte Carlo sampling. We prove a result that characterizes optimal e-values for group invariance against optimality targets that satisfy a mild orbit-wise decomposition property. We apply this to design expected-utility-optimal e-values for group invariance, which include both Neyman-Pearson-optimal tests and log-optimal e-values. Moreover, we generalize the notion of rank- and sign-based testing to compact groups, by using a representative inversion kernel. In addition, we characterize e-processes for group invariance for arbitrary filtrations, and provide tools to construct them. We also describe test martingales under a natural filtration, which are simpler to construct. Peeking beyond compact groups, we encounter e-values and e-processes based on ergodic theorems. These nest e-processes based on de Finetti's theorem for testing exchangeability.

math.ST

Anytime Validity is Free: Inducing Sequential Tests

Anytime valid sequential tests permit us to stop testing based on the current data, without invalidating the inference. Given a maximum number of observations $N$, one may believe this must come at the cost of power when compared to a conventional test that waits until all $N$ observations have arrived. Our first contribution is to show that this is false: for any valid test based on $N$ observations, we show how to construct an anytime valid sequential test that matches it after $N$ observations. Our second contribution is that we may continue testing by using the outcome of a $[0, 1]$-valued test as a conditional significance level in subsequent testing, leading to an overall procedure that is valid at the original significance level. This shows that anytime validity and optional continuation are readily available in traditional testing, without requiring explicit use of e-values. We illustrate this by deriving the anytime valid sequentialized $z$-test and $t$-test, which at time $N$ coincide with the traditional $z$-test and $t$-test. Finally, we characterize the SPRT by invariance under test induction, and also show under an i.i.d. assumption that the SPRT is induced by the Neyman-Pearson test for a tiny significance level and huge $N$.

math.ST

Post-hoc $α$ Hypothesis Testing and the Post-hoc $p$-value

In traditional hypothesis testing one must pre-specify the significance level $α$ to bound the `size' of the test: its probability to falsely reject the hypothesis. Indeed, a data-dependent selection of $α$ would generally distort the size, possibly making it larger than the specified level $α$. We explore hypothesis testing with a data-dependent choice of $α$ by guaranteeing that there is no such size distortion in expectation, even if the level $α$ is arbitrarily selected based on the data. Unlike regular $p$-values, resulting `post-hoc $p$-values' allow us to `reject at level $p$' and still provide this guarantee. Interestingly, we find that $p$ is a post-hoc $p$-value if and only if $1/p$ is an $e$-value, a recently introduced measure of evidence. While often treated as different paradigms, this reveals $e$-values are simply $p$-values under a stronger error guarantee, thinly veiled by the reciprocal $p = 1/e$. Moreover, we extend classical optimal testing to optimal post-hoc testing. Finally, we apply our work to close Markov's inequality into a post-hoc $α$ equality, and we study more general forms of post-hoc testing that require us to generalize beyond $e$-values.

math.ST

Assessing solution quality in risk-averse stochastic programs

In optimization problems, the quality of a candidate solution can be characterized by the optimality gap. For most stochastic optimization problems, this gap must be statistically estimated. We show that for risk-averse problems, standard estimators are optimistically biased, which compromises the statistical guarantee on the optimality gap. We introduce estimators for risk-averse problems that do not suffer from this bias. Our method relies on using two independent samples, each estimating a different component of the optimality gap. Our approach extends a broad class of optimality gap estimation methods from the risk-neutral case to the risk-averse case, such as the multiple replications procedure and its one- and two-sample variants. We show that our approach is tractable and leads to high-quality optimality gap estimates for spectral and quadrangle risk measures. Our approach can further make use of existing bias and variance reduction techniques.

math.OC

Real-time Program Evaluation using Anytime-valid Rank Tests

Counterfactual mean estimators such as difference-in-differences and synthetic control have grown into workhorse tools for program evaluation. Inference for these estimators is well-developed in settings where all post-treatment data is available at the time of analysis. However, in settings where data arrives sequentially, these tests do not permit real-time inference, as they require a pre-specified sample size T. We introduce real-time inference for program evaluation through anytime-valid rank tests. Our methodology relies on interpreting the absence of a treatment effect as exchangeability of the treatment estimates. We then convert these treatment estimates into sequential ranks, and construct optimal finite-sample valid sequential tests for exchangeability. We illustrate our methods in the context of difference-in-differences and synthetic control. In simulations, they control size even under mild exchangeability violations. While our methods suffer slight power loss at T, they allow for early rejection (before T) and preserve the ability to reject later (after T).

econ.EM

More Power by using Fewer Permutations

It is conventionally believed that a permutation test should ideally use all permutations. If this is computationally unaffordable, it is believed one should use the largest affordable Monte Carlo sample or (algebraic) subgroup of permutations. We challenge this belief by showing we can sometimes obtain dramatically more power by using a tiny subgroup. As the subgroup is tiny, this simultaneously comes at a much lower computational cost. We exploit this to improve the popular permutation-based Westfall & Young MaxT multiple testing method. We study the relative efficiency in a Gaussian location model, and find the largest gain in high dimensions.

math.ST

More Efficient Exact Group-Invariance Testing: using a Representative Subgroup

Non-parametric tests based on permutation, rotation or sign-flipping are examples of group-invariance tests. These tests test invariance of the null distribution under a set of transformations that has a group structure, in the algebraic sense. Such groups are often huge, which makes it computationally infeasible to test using the entire group. Hence, it is standard practice to test using a randomly sampled set of transformations from the group. This random sample still needs to be substantial to obtain good power and replicability. We improve upon this standard practice by using a well-designed subgroup of transformations instead of a random sample. The resulting subgroup-invariance test is still exact, as invariance under a group implies invariance under its subgroups. We illustrate this in a generalized location model and obtain more powerful tests based on the same number of transformations. In particular, we show that a subgroup-invariance test is consistent for lower signal-to-noise ratios than a test based on a random sample. For the special case of a normal location model and a particular design of the subgroup, we show that the power improvement is equivalent to the power difference between a Monte Carlo $Z$-test and a Monte Carlo $t$-test.

stat.ME