Searcharxiv⌕ Search

arXiv subjects

Abba M. Krieger

Publications and source records attributed to Abba M. Krieger.

13 recordsLinked to original sources

Optimal Designs with Robust Inference for Binary Treatment Effects

We study randomized experiments with binary outcomes under Neyman's nonparametric model, where covariate measurements are fixed but potential outcomes are random. In this setting we derive the exact variance of the difference-in-means estimator and characterize designs that minimize it. We show that any balanced design satisfying a covariate-balance condition is asymptotically optimal, and we prove that a broad class of blocking designs satisfies this condition under mild smoothness assumptions. Because the variance depends on unknown success probabilities, unbiased variance estimation is impossible. We therefore develop two conservative estimators: a generalization of the Cochran-Mantel-Haenszel (CMH) statistic applicable to any balanced design, and an extension of Robins' variance estimator for blocking designs. We establish conditions under which the CMH-based estimator is asymptotically tight under local alternatives, thereby yielding asymptotically valid confidence intervals. Our theoretical and simulations results show that blocking and other designs that achieve covariate balance and sufficient degree of randomness, perform well when they are equipped with the CMH-based inference. Thus, we provide an experimental design and inference framework for incidence outcomes that is simultaneously variance-optimal, conservative in finite samples and asymptotically tight.

stat.ME↗

Block Designs that Provide Optimal Power in the Cochran-Mantel-Haenszel Test

We consider the asymptotic power performance under local alternatives of the Cochran-Mantel-Haenszel test. Our setting is non-traditional: we investigate randomized experiments that assign subjects via Fisher's blocking design. We show that blocking designs that satisfy a certain balance condition are asymptotically optimal. When the potential outcomes can be ordered, the balance condition is met for all blocking designs with number of blocks going to infinity. More generally, we prove that the pairwise matching design of Greevy et al. (2004) satisfies the balance condition under mild assumptions. In smaller sample sizes, we show a second order effect becomes operational thereby making blocking designs with a smaller number optimal. In practical settings with many covariates, we recommend pairwise matching for its ability to approximate the balance condition.

stat.ME↗

The Optimality of Blocking Designs in Equally and Unequally Allocated Randomized Experiments with General Response

We consider the performance of the difference-in-means estimator in a two-arm randomized experiment under common experimental endpoints such as continuous (regression), incidence, proportion and survival. We examine performance under both equal and unequal allocation to treatment groups and we consider both the Neyman randomization model and the population model. We show that in the Neyman model, where the only source of randomness is the treatment manipulation, there is no free lunch: complete randomization is minimax for the estimator's mean squared error. In the population model, where each subject experiences response noise with zero mean, the optimal design is the deterministic perfect-balance allocation. However, this allocation is generally NP-hard to compute and moreover, depends on unknown response parameters. When considering the tail criterion of Kapelner et al. (2021), we show the optimal design is less random than complete randomization and more random than the deterministic perfect-balance allocation. We prove that Fisher's blocking design provides the asymptotically optimal degree of experimental randomness. Theoretical results are supported by simulations in all considered experimental settings.

math.ST↗

The Pairwise Matching Design is Optimal under Extreme Noise and Assignments

We consider the general performance of the difference-in-means estimator in an equally-allocated two-arm randomized experiment under common experimental endpoints such as continuous (regression), incidence, proportion, count and uncensored survival. We consider two sources of randomness: the subject-specific assignments and the contribution of unobserved subject-specific measurements. We then examine mean squared error (MSE) performance under a new, more realistic "simultaneous tail criterion". We prove that the pairwise matching design of Greevy et al. (2004) performs best asymptotically under this criterion when compared to other blocking designs. We also prove that the optimal design must be less random than complete randomization and more random than any deterministic, optimized allocation. Theoretical results are supported by simulations in all five response types.

stat.ME↗

The Role of Pairwise Matching in Experimental Design for an Incidence Outcome

We consider the problem of evaluating designs for a two-arm randomized experiment with an incidence (binary) outcome under a nonparametric general response model. Our two main results are that the priori pair matching design of Greevy et al. (2004) is (1) the optimal design as measured by mean squared error among all block designs which includes complete randomization. And (2), this pair-matching design is minimax, i.e. it provides the lowest mean squared error under an adversarial response model. Theoretical results are supported by simulations and clinical trial data.

stat.ME↗

Better Experimental Design by Hybridizing Binary Matching with Imbalance Optimization

We present a new experimental design procedure that divides a set of experimental units into two groups in order to minimize error in estimating an additive treatment effect. One concern is minimizing error at the experimental design stage is large covariate imbalance between the two groups. Another concern is robustness of design to misspecification in response models. We address both concerns in our proposed design: we first place subjects into pairs using optimal nonbipartite matching, making our estimator robust to complicated non-linear response models. Our innovation is to keep the matched pairs extant, take differences of the covariate values within each matched pair and then we use the greedy switching heuristic of Krieger et al. (2019) or rerandomization on these differences. This latter step greatly reduce covariate imbalance to the rate $O_p(n^{-4})$ in the case of one covariate that are uniformly distributed. This rate benefits from the greedy switching heuristic which is $O_p(n^{-3})$ and the rate of matching which is $O_p(n^{-1})$. Further, our resultant designs are shown to be as random as matching which is robust to unobserved covariates. When compared to previous designs, our approach exhibits significant improvement in the mean squared error of the treatment effect estimator when the response model is nonlinear and performs at least as well when it the response model is linear. Our design procedure is found as a method in the open source R package available on CRAN called GreedyExperimentalDesign.

stat.ME↗

Optimal Rerandomization via a Criterion that Provides Insurance Against Failed Experiments

We present an optimized rerandomization design procedure for a non-sequential treatment-control experiment. Randomized experiments are the gold standard for finding causal effects in nature. But sometimes random assignments result in unequal partitions of the treatment and control group visibly seen as imbalance in observed covariates. There can additionally be imbalance on unobserved covariates. Imbalance in either observed or unobserved covariates increases treatment effect estimator error inflating the width of confidence regions and reducing experimental power. "Rerandomization" is a strategy that omits poor imbalance assignments by limiting imbalance in the observed covariates to a prespecified threshold. However, limiting this threshold too much can increase the risk of contracting error from unobserved covariates. We introduce a criterion that combines observed imbalance while factoring in the risk of inadvertently imbalancing unobserved covariates. We then use this criterion to locate the optimal rerandomization threshold based on the practitioner's level of desired insurance against high estimator error. We demonstrate the gains of our designs in simulation and in a dataset from a large randomized experiment in education. We provide an open source R package available on CRAN named OptimalRerandExpDesigns which generates designs according to our algorithm.

stat.ME↗

Improving the Power of the Randomization Test

We consider the problem of evaluating designs for a two-arm randomized experiment with the criterion being the power of the randomization test for the one-sided null hypothesis. Our evaluation assumes a response that is linear in one observed covariate, an unobserved component and an additive treatment effect where the only randomness comes from the treatment allocations. It is well-known that the power depends on the allocations' imbalance in the observed covariate and this is the reason for the classic restricted designs such as rerandomization. We show that power is also affected by two other design choices: the number of allocations in the design and the degree of linear dependence among the allocations. We prove that the more allocations, the higher the power and the lower the variability in the power. Designs that feature greater independence of allocations are also shown to have higher performance. Our theoretical findings and extensive simulation studies imply that the designs with the highest power provide thousands of highly independent allocations that each provide nominal imbalance in the observed covariates. These high powered designs exhibit less randomization than complete randomization and more randomization than recently proposed designs based on numerical optimization. Model choices for a practicing experimenter are rerandomization and greedy pair switching, where both outperform complete randomization and numerical optimization. The tradeoff we find also provides a means to specify the imbalance threshold parameter when rerandomizing.

stat.ME↗

Harmonizing Fully Optimal Designs with Classic Randomization in Fixed Trial Experiments

There is a movement in design of experiments away from the classic randomization put forward by Fisher, Cochran and others to one based on optimization. In fixed-sample trials comparing two groups, measurements of subjects are known in advance and subjects can be divided optimally into two groups based on a criterion of homogeneity or "imbalance" between the two groups. These designs are far from random. This paper seeks to understand the benefits and the costs over classic randomization in the context of different performance criterions such as Efron's worst-case analysis. In the criterion that we motivate, randomization beats optimization. However, the optimal design is shown to lie between these two extremes. Much-needed further work will provide a procedure to find this optimal designs in different scenarios in practice. Until then, it is best to randomize.

stat.ME↗

Nearly Random Designs with Greatly Improved Balance

We present a new experimental design procedure that divides a set of experimental units into two groups so that the two groups are balanced on a prespecified set of covariates and being almost as random as complete randomization. Under complete randomization, the difference in covariate balance as measured by the standardized average between treatment and control will be $O_p(n^{-1/2})$. If the sample size is not too large this may be material. In this article, we present an algorithm which greedily switches assignment pairs. Resultant designs produce balance of the much lower order $O_p(n^{-3})$ for one covariate. However, our algorithm creates assignments which are, strictly speaking, non-random. We introduce two metrics which capture departures from randomization: one in the style of entropy and one in the style of standard error and demonstrate our assignments are nearly as random as complete randomization in terms of both measures. The results are extended to more than one covariate, simulations are provided to illustrate the results and statistical inference under our design is discussed. We provide an open source R package available on CRAN called GreedyExperimentalDesign which generates designs according to our algorithm.

math.ST↗

Extreme(ly) mean(ingful): Sequential formation of a quality group

The present paper studies the limiting behavior of the average score of a sequentially selected group of items or individuals, the underlying distribution of which, $F$, belongs to the Gumbel domain of attraction of extreme value distributions. This class contains the Normal, Lognormal, Gamma, Weibull and many other distributions. The selection rules are the "better than average" ($β=1$) and the "$β$-better than average" rule, defined as follows. After the first item is selected, another item is admitted into the group if and only if its score is greater than $β$ times the average score of those already selected. Denote by $\bar{Y}_k$ the average of the $k$ first selected items, and by $T_k$ the time it takes to amass them. Some of the key results obtained are: under mild conditions, for the better than average rule, $\bar{Y}_k$ less a suitable chosen function of $\log k$ converges almost surely to a finite random variable. When $1-F(x)=e^{-[x^α+h(x)]}$, $α>0$ and $h(x)/x^α\stackrel{x\rightarrow \infty}{\longrightarrow}0$, then $T_k$ is of approximate order $k^2$. When $β>1$, the asymptotic results for $\bar{Y}_k$ are of a completely different order of magnitude. Interestingly, for a class of distributions, $T_k$, suitably normalized, asymptotically approaches 1, almost surely for relatively small $β\ge1$, in probability for moderate sized $β$ and in distribution when $β$ is large.

math.PR↗

Select sets: Rank and file

In many situations, the decision maker observes items in sequence and needs to determine whether or not to retain a particular item immediately after it is observed. Any decision rule creates a set of items that are selected. We consider situations where the available information is the rank of a present observation relative to its predecessors. Certain ``natural'' selection rules are investigated. Theoretical results are presented pertaining to the evolution of the number of items selected, measures of their quality and the time it would take to amass a group of a given size.

math.PR↗

R-Estimates vs. GMM: A Theoretical Case Study of Validity and Efficiency

What role should assumptions play in inference? We present a small theoretical case study of a simple, clean case, namely the nonparametric comparison of two continuous distributions using (essentially) information about quartiles, that is, the central information displayed in a pair of boxplots. In particular, we contrast a suggestion of John Tukey--that the validity of inferences should not depend on assumptions, but assumptions have a role in efficiency--with a competing suggestion that is an aspect of Hansen's generalized method of moments--that methods should achieve maximum asymptotic efficiency with fewer assumptions. In our case study, the practical performance of these two suggestions is strikingly different. An aspect of this comparison concerns the unification or separation of the tasks of estimation assuming a model and testing the fit of that model. We also look at a method (MERT) that aims not at best performance, but rather at achieving reasonable performance across a set of plausible models.

math.ST↗