SearcharxivSearch

arXiv subjects

David Preinerstorfer

Publications and source records attributed to David Preinerstorfer.

At least 19 recordsLinked to original sources

Adversarially robust multiple testing in high dimensions

Robust multiple testing procedures for assessing equality restrictions on the coordinates of high-dimensional mean vectors are proposed. Our procedures are based on quantile-winsorization techniques, approximately control the familywise error rate (strongly) under adversarial contamination, and allow the number of hypotheses to grow exponentially with sample size, despite requiring only slightly more than two moments. Technically, our one-sample results build on recent Gaussian approximation inequalities for the distribution of high-dimensional quantile-winsorized means, whereas our two-sample results are based on extensions thereof, which we develop here and which could be of some interest in their own right.

math.ST

Robust Instrumental Variables: Sharp Rates and Inference under Adversarial Contamination

Because 2SLS is built from sample averages, a small number of observations can have a disproportionate effect on estimates and inference. We introduce W-2SLS, a simple drop-in robustification that replaces these averages by quantile-winsorized means. We analyze W-2SLS under adversarial contamination, which permits both the identities and the reported values of the contaminated observations to depend on the realized clean sample and therefore accommodates targeted or strategic manipulation. Under finite $m$-th moments, W-2SLS attains the minimax-sharp rate $\eta_{n}^{1-\frac1m}+n^{-1/2}$, where $\eta_n$ is the fraction of observations that may be altered. Matching lower bounds identify the exact contamination thresholds for uniform consistency, root-$n$ estimation, and centered Gaussian inference with the same first-order law as clean-sample 2SLS. When $\sqrt{n}\eta_{n}^{1-\frac1m}\to 0$ robustness is first-order free. We also construct feasible heteroskedasticity-robust inference and a winsorized Anderson--Rubin test valid under weak identification and adversarial contamination. Finally, even without contamination, ordinary 2SLS can have poor uniform finite-sample concentration, whereas W-2SLS admits confidence-calibrated sub-Gaussian deviation guarantees.

econ.EM

Robustness for free: asymptotic size and power of max-tests in high dimensions

Allowing for adversarial contamination and heavy tails, we study testing whether the mean of a high-dimensional random vector equals zero. Because standard max-tests based on sample averages are highly non-robust, we propose a max-test based on quantile-winsorized observations. The test controls asymptotic size under adversarial contamination and only requires $m>2$ moments, while allowing dimension to grow exponentially with sample size. We fully characterize its asymptotic power function. Comparing with the standard max-test, for which we also derive a power characterization as a benchmark, we show that robustness is obtained for free: under the stronger conditions that make the standard max-test valid, our robust test has identical asymptotic power. We also study the role of bootstrap critical values, showing that their use never decreases power, can strictly improve asymptotic power in extremely correlated designs, but often has no first-order asymptotic effect.

math.ST

High-dimensional Gaussian and bootstrap approximations for robust means

Recent years have witnessed much progress on Gaussian and bootstrap approximations to the distribution of sums of independent random vectors with dimension $d$ large relative to the sample size $n$. However, for any number of moments $m>2$ that the summands may possess, there exist distributions such that these approximations break down if $d$ grows faster than the polynomial barrier $n^{\frac{m}{2}-1}$. In this paper, we establish Gaussian and bootstrap approximations to the distributions of winsorized and trimmed means that allow $d$ to grow at an exponential rate in $n$ as long as $m>2$ moments exist. The approximations remain valid under some amount of adversarial contamination. Our implementations of the winsorized and trimmed means do not require knowledge of $m$. As a consequence, the approximation guarantees ``adapt'' to $m$.

math.ST

Winsorized mean estimation with heavy tails and adversarial contamination

Finite-sample upper bounds on the estimation error of a winsorized mean estimator of the population mean in the presence of heavy tails and adversarial contamination are established. In comparison to existing results, the winsorized mean estimator we study avoids a sample-splitting device and winsorizes substantially fewer observations, which improves its applicability and practical performance.

math.ST

A Necessary and Sufficient Condition for Size Controllability of Heteroskedasticity Robust Test Statistics

We revisit size controllability results in P\"otscher and Preinerstorfer (2025) concerning heteroskedasticity robust test statistics in regression models. For the special, but important, case of testing a single restriction (e.g., a zero restriction on a single coefficient), we povide a necessary and sufficient condition for size controllability, whereas the condition in P\"otscher and Preinerstorfer (2025) is, in general, only sufficient (even in the case of testing a single restriction).

math.ST

Enhanced power enhancements for testing many moment equalities: Beyond the $2$- and $\infty$-norm

Contemporary testing problems in statistics are increasingly complex, i.e., high-dimensional. Tests based on the $2$- and $\infty$-norm have received considerable attention in such settings, as they are powerful against dense and sparse alternatives, respectively. The power enhancement principle of Fan et al. (2015) combines these two norms to construct improved tests that are powerful against both types of alternatives. In the context of testing whether a candidate parameter satisfies a large number of moment equalities, we construct a test that harnesses the strength of all $p$-norms with $p\in[2, \infty]$. As a result, this test is consistent against strictly more alternatives than any test based on a single $p$-norm. In particular, our test is consistent against more alternatives than tests based on the $2$- and $\infty$-norm, which is what most implementations of the power enhancement principle target. We illustrate the scope of our general results by using them to construct a test that simultaneously dominates the Anderson-Rubin test (based on $p=2$), tests based on the $\infty$-norm and power enhancement based combinations of these in terms of consistency in the linear instrumental variable model with many instruments.

econ.EM

Regularizing Fairness in Optimal Policy Learning with Distributional Targets

A decision maker typically (i) incorporates training data to learn about the relative effectiveness of treatments, and (ii) chooses an implementation mechanism that implies an ``optimal'' predicted outcome distribution according to some target functional. Nevertheless, a fairness-aware decision maker may not be satisfied achieving said optimality at the cost of being ``unfair" against a subgroup of the population, in the sense that the outcome distribution in that subgroup deviates too strongly from the overall optimal outcome distribution. We study a framework that allows the decision maker to regularize such deviations, while allowing for a wide range of target functionals and fairness measures to be employed. We establish regret and consistency guarantees for empirical success policies with (possibly) data-driven preference parameters, and provide numerical results. Furthermore, we briefly illustrate the methods in two empirical settings.

econ.EM

A Modern Gauss-Markov Theorem? Really?

We show that the theorems in Hansen (2021a) (the version accepted by Econometrica), except for one, are not new as they coincide with classical theorems like the good old Gauss-Markov or Aitken Theorem, respectively; the exceptional theorem is incorrect. Hansen (2021b) corrects this theorem. As a result, all theorems in the latter version coincide with the above mentioned classical theorems. Furthermore, we also show that the theorems in Hansen (2022) (the version published in Econometrica) either coincide with the classical theorems just mentioned, or contain extra assumptions that are alien to the Gauss-Markov or Aitken Theorem.

math.ST

Treatment recommendation with distributional targets

We study the problem of a decision maker who must provide the best possible treatment recommendation based on an experiment. The desirability of the outcome distribution resulting from the policy recommendation is measured through a functional capturing the distributional characteristic that the decision maker is interested in optimizing. This could be, e.g., its inherent inequality, welfare, level of poverty or its distance to a desired outcome distribution. If the functional of interest is not quasi-convex or if there are constraints, the optimal recommendation may be a mixture of treatments. This vastly expands the set of recommendations that must be considered. We characterize the difficulty of the problem by obtaining maximal expected regret lower bounds. Furthermore, we propose two (near) regret-optimal policies. The first policy is static and thus applicable irrespectively of subjects arriving sequentially or not in the course of the experimentation phase. The second policy can utilize that subjects arrive sequentially by successively eliminating inferior treatments and thus spends the sampling effort where it is most needed.

econ.EM

Consistency of $p$-norm based tests in high dimensions: characterization, monotonicity, domination

Many commonly used test statistics are based on a norm measuring the evidence against the null hypothesis. To understand how the choice of a norm affects power properties of tests in high dimensions, we study the consistency sets of $p$-norm based tests in the prototypical framework of sequence models with unrestricted parameter spaces, the null hypothesis being that all observations have zero mean. The consistency set of a test is here defined as the set of all arrays of alternatives the test is consistent against as the dimension of the parameter space diverges. We characterize the consistency sets of $p$-norm based tests and find, in particular, that the consistency against an array of alternatives cannot be determined solely in terms of the $p$-norm of the alternative. Our characterization also reveals an unexpected monotonicity result: namely that the consistency set is strictly increasing in $p \in (0, \infty)$, such that tests based on higher $p$ strictly dominate those based on lower $p$ in terms of consistency. This monotonicity allows us to construct novel tests that dominate, with respect to their consistency behavior, all $p$-norm based tests without sacrificing size.

math.ST

Superconsistency of Tests in High Dimensions

To assess whether there is some signal in a big database, aggregate tests for the global null hypothesis of no effect are routinely applied in practice before more specialized analysis is carried out. Although a plethora of aggregate tests is available, each test has its strengths but also its blind spots. In a Gaussian sequence model, we study whether it is possible to obtain a test with substantially better consistency properties than the likelihood ratio (i.e., Euclidean norm based) test. We establish an impossibility result, showing that in the high-dimensional framework we consider, the set of alternatives for which a test may improve upon the likelihood ratio test -- that is, its superconsistency points -- is always asymptotically negligible in a relative volume sense.

math.ST

How Reliable are Bootstrap-based Heteroskedasticity Robust Tests?

We develop theoretical finite-sample results concerning the size of wild bootstrap-based heteroskedasticity robust tests in linear regression models. In particular, these results provide an efficient diagnostic check, which can be used to weed out tests that are unreliable for a given testing problem in the sense that they overreject substantially. This allows us to assess the reliability of a large variety of wild bootstrap-based tests in an extensive numerical study.

math.ST

Valid Heteroskedasticity Robust Testing

Tests based on heteroskedasticity robust standard errors are an important technique in econometric practice. Choosing the right critical value, however, is not simple at all: conventional critical values based on asymptotics often lead to severe size distortions; and so do existing adjustments including the bootstrap. To avoid these issues, we suggest to use smallest size-controlling critical values, the generic existence of which we prove in this article for the commonly used test statistics. Furthermore, sufficient and often also necessary conditions for their existence are given that are easy to check. Granted their existence, these critical values are the canonical choice: larger critical values result in unnecessary power loss, whereas smaller critical values lead to over-rejections under the null hypothesis, make spurious discoveries more likely, and thus are invalid. We suggest algorithms to numerically determine the proposed critical values and provide implementations in accompanying software. Finally, we numerically study the behavior of the proposed testing procedures, including their power properties.

math.ST

Functional Sequential Treatment Allocation

Consider a setting in which a policy maker assigns subjects to treatments, observing each outcome before the next subject arrives. Initially, it is unknown which treatment is best, but the sequential nature of the problem permits learning about the effectiveness of the treatments. While the multi-armed-bandit literature has shed much light on the situation when the policy maker compares the effectiveness of the treatments through their mean, much less is known about other targets. This is restrictive, because a cautious decision maker may prefer to target a robust location measure such as a quantile or a trimmed mean. Furthermore, socio-economic decision making often requires targeting purpose specific characteristics of the outcome distribution, such as its inherent degree of inequality, welfare or poverty. In the present paper we introduce and study sequential learning algorithms when the distributional characteristic of interest is a general functional of the outcome distribution. Minimax expected regret optimality results are obtained within the subclass of explore-then-commit policies, and for the unrestricted class of all policies.

econ.EM

Functional Sequential Treatment Allocation with Covariates

We consider a multi-armed bandit problem with covariates. Given a realization of the covariate vector, instead of targeting the treatment with highest conditional expectation, the decision maker targets the treatment which maximizes a general functional of the conditional potential outcome distribution, e.g., a conditional quantile, trimmed mean, or a socio-economic functional such as an inequality, welfare or poverty measure. We develop expected regret lower bounds for this problem, and construct a near minimax optimal assignment policy.

stat.ML

Further Results on Size and Power of Heteroskedasticity and Autocorrelation Robust Tests, with an Application to Trend Testing

We complement the theory developed in Preinerstorfer and Pötscher (2016) with further finite sample results on size and power of heteroskedasticity and autocorrelation robust tests. These allows us, in particular, to show that the sufficient conditions for the existence of size-controlling critical values recently obtained in Pötscher and Preinerstorfer (2018) are often also necessary. We furthermore apply the results obtained to tests for hypotheses on deterministic trends in stationary time series regressions, and find that many tests currently used are strongly size-distorted.

math.ST