SearcharxivSearch

arXiv subjects

Joseph P. Romano

Publications and source records attributed to Joseph P. Romano.

At least 19 recordsLinked to original sources

Least Squares-Based Permutation Tests in Time Series

This paper studies permutation tests for regression parameters in a time series setting, where the time series is assumed stationary but may exhibit an arbitrary (but weak) dependence structure. In such a setting, it is perhaps surprising that permutation tests can offer any type of inference guarantees, since permuting of covariates can destroy their relationship with the response. Indeed, the fundamental assumption of exchangeability of errors required for the finite-sample exactness of permutation tests can easily fail. However, we show that permutation tests may be constructed which are asymptotically valid for a wide class of stationary processes, but remain exact when exchangeability holds. We also consider the problem of testing for no monotone trend and we construct asymptotically valid permutation tests in this setting as well. In addition, in order to use the methods, the R package permixOLS is publicly available.

math.ST

Graph-Laplacian Variance Estimators for Finely Stratified Experiments

This paper considers design-based inference on the average treatment effect in finely stratified experiments, where uncertainty arises only from the randomized treatment assignment. We focus on settings in which units are first stratified into groups of fixed size according to baseline covariates and, then within each group, exactly one unit is assigned to treatment. In this setting, we introduce a class of graph-Laplacian variance estimators in which strata form the vertices of a weighted graph and edge weights determine how between-stratum comparisons are aggregated. The canonical estimator of Imai (2008) corresponds to a complete graph with edge weights normalized so that each stratum has weighted degree one, while a paired-stratum estimator arises from a perfect matching graph. For the subclass of degree-calibrated graphs, in which each vertex has weighted degree one, we derive an exact bias identity showing that the corresponding estimators are upward-biased, with bias governed by squared differences in the true stratum-level treatment effects across adjacent strata. As a result, any such estimator may be used for valid inference. The identity further suggests that paired-stratum estimators constructed from a covariate-based perfect matching can induce small biases when treatment effects vary smoothly with the covariates. Without such smoothness, however, we show that paired-stratum estimators can exhibit large worst-case bias, and that, within the class of degree-calibrated estimators, the complete-graph estimator is minimax optimal for normalized bias under a weak bound on treatment-effect heterogeneity. Motivated by this contrast, we propose a regularized graph estimator that controls worst-case normalized bias while preserving much of the locality of the paired-stratum estimator. Simulations illustrate the resulting tradeoff between locality and worst-case protection.

econ.EM

A New Design-Based Variance Estimator for Finely Stratified Experiments

This paper considers the problem of design-based inference for the average treatment effect in finely stratified experiments. Here, by "design-based'' we mean that the only source of uncertainty stems from the randomness in treatment assignment; by "finely stratified'' we mean that units are stratified into groups of a fixed size according to baseline covariates and then, within each group, a fixed number of units are assigned uniformly at random to treatment and the remainder to control. In this setting we present a novel estimator of the variance of the difference-in-means based on pairing "adjacent" strata. Importantly, our estimator is well defined even in the challenging setting where there is exactly one treated or control unit per stratum. We prove that our estimator is upward-biased, and thus can be used for inference under mild restrictions on the finite population. We compare our estimator with some well-known estimators that have been proposed previously in this setting, and demonstrate that, while these estimators are also upward-biased, our estimator has smaller bias and therefore leads to more precise inferences whenever adjacent strata are sufficiently similar. To further understand when our estimator leads to more precise inferences, we introduce a framework motivated by a thought experiment in which the finite population is modeled as having been drawn once in an i.i.d. fashion from a well-behaved probability distribution. In this framework, we argue that our estimator dominates the others in terms of limiting bias and that these improvements are strict except under strong restrictions on the treatment effects. Finally, we illustrate the practical relevance of our theoretical results through a simulation study, which reveals that our estimator can in fact lead to substantially more precise inferences, especially when the quality of stratification is high.

econ.EM

Randomization Inference: Theory and Applications

We review approaches to statistical inference based on randomization. Permutation tests are treated as an important special case. Under a certain group invariance property, referred to as the ``randomization hypothesis,'' randomization tests achieve exact control of the Type I error rate in finite samples. Although this unequivocal precision is very appealing, the range of problems that satisfy the randomization hypothesis is somewhat limited. We show that randomization tests are often asymptotically, or approximately, valid and efficient in settings that deviate from the conditions required for finite-sample error control. When randomization tests fail to offer even asymptotic Type 1 error control, their asymptotic validity may be restored by constructing an asymptotically pivotal test statistic. Randomization tests can then provide exact error control for tests of highly structured hypotheses with good performance in a wider class of problems. We give a detailed overview of several prominent applications of randomization tests, including two-sample permutation tests, regression, and conformal inference.

econ.EM

Reproducible Aggregation of Sample-Split Statistics

Statistical inference is often simplified by sample-splitting. This simplification comes at the cost of the introduction of randomness not native to the data. We propose a simple procedure for sequentially aggregating statistics constructed with multiple splits of the same sample. The user specifies a bound and a nominal error rate. If the procedure is implemented twice on the same data, the nominal error rate approximates the chance that the results differ by more than the bound. We illustrate the application of the procedure to several widely applied econometric methods.

econ.EM

Permutation Testing for Monotone Trend

In this paper, we consider the fundamental problem of testing for monotone trend in a time series. While the term "trend" is commonly used and has an intuitive meaning, it is first crucial to specify its exact meaning in a hypothesis testing context. A commonly used well-known test is the Mann-Kendall test, which we show does not offer Type 1 error control even in large samples. On the other hand, by an appropriate studentization of the Mann-Kendall statistic, we construct permutation tests that offer asymptotic error control quite generally, but retain the exactness property of permutation tests for i.i.d. observations. We also introduce "local" Mann-Kendall statistics as a means of testing for local rather than global trend in a time series. Similar properties of permutation tests are obtained for these tests as well.

math.ST

Covariate Adjustment in Experiments with Matched Pairs

This paper studies inference on the average treatment effect in experiments in which treatment status is determined according to "matched pairs" and it is additionally desired to adjust for observed, baseline covariates to gain further precision. By a "matched pairs" design, we mean that units are sampled i.i.d. from the population of interest, paired according to observed, baseline covariates and finally, within each pair, one unit is selected at random for treatment. Importantly, we presume that not all observed, baseline covariates are used in determining treatment assignment. We study a broad class of estimators based on a "doubly robust" moment condition that permits us to study estimators with both finite-dimensional and high-dimensional forms of covariate adjustment. We find that estimators with finite-dimensional, linear adjustments need not lead to improvements in precision relative to the unadjusted difference-in-means estimator. This phenomenon persists even if the adjustments are interacted with treatment; in fact, doing so leads to no changes in precision. However, gains in precision can be ensured by including fixed effects for each of the pairs. Indeed, we show that this adjustment is the "optimal" finite-dimensional, linear adjustment. We additionally study two estimators with high-dimensional forms of covariate adjustment based on the LASSO. For each such estimator, we show that it leads to improvements in precision relative to the unadjusted difference-in-means estimator and also provide conditions under which it leads to the "optimal" nonparametric, covariate adjustment. A simulation study confirms the practical relevance of our theoretical analysis, and the methods are employed to reanalyze data from an experiment using a "matched pairs" design to study the effect of macroinsurance on microenterprise.

econ.EM

Confidence Intervals for Seroprevalence

This paper concerns the construction of confidence intervals in standard seroprevalence surveys. In particular, we discuss methods for constructing confidence intervals for the proportion of individuals in a population infected with a disease using a sample of antibody test results and measurements of the test's false positive and false negative rates. We begin by documenting erratic behavior in the coverage probabilities of standard Wald and percentile bootstrap intervals when applied to this problem. We then consider two alternative sets of intervals constructed with test inversion. The first set of intervals are approximate, using either asymptotic or bootstrap approximation to the finite-sample distribution of a chosen test statistic. We consider several choices of test statistic, including maximum likelihood estimators and generalized likelihood ratio statistics. We show with simulation that, at empirically relevant parameter values and sample sizes, the coverage probabilities for these intervals are close to their nominal level and are approximately equi-tailed. The second set of intervals are shown to contain the true parameter value with probability at least equal to the nominal level, but can be conservative in finite samples.

stat.AP

Uncertainty in the Hot Hand Fallacy: Detecting Streaky Alternatives to Random Bernoulli Sequences

We study a class of permutation tests of the randomness of a collection of Bernoulli sequences and their application to analyses of the human tendency to perceive streaks of consecutive successes as overly representative of positive dependence - the hot hand fallacy. In particular, we study permutation tests of the null hypothesis of randomness (i.e., that trials are i.i.d.) based on test statistics that compare the proportion of successes that directly follow k consecutive successes with either the overall proportion of successes or the proportion of successes that directly follow k consecutive failures. We characterize the asymptotic distributions of these test statistics and their permutation distributions under randomness, under a set of general stationary processes, and under a class of Markov chain alternatives, which allow us to derive their local asymptotic power. The results are applied to evaluate the empirical support for the hot hand fallacy provided by four controlled basketball shooting experiments. We establish that substantially larger data sets are required to derive an informative measurement of the deviation from randomness in basketball shooting. In one experiment, for which we were able to obtain data, multiple testing procedures reveal that one shooter exhibits a shooting pattern significantly inconsistent with randomness - supplying strong evidence that basketball shooting is not random for all shooters all of the time. However, we find that the evidence against randomness in this experiment is limited to this shooter. Our results provide a mathematical and statistical foundation for the design and validation of experiments that directly compare deviations from randomness with human beliefs about deviations from randomness, and thereby constitute a direct test of the hot hand fallacy.

econ.EM

Permutation Testing for Dependence in Time Series

Given observations from a stationary time series, permutation tests allow one to construct exactly level $α$ tests under the null hypothesis of an i.i.d. (or, more generally, exchangeable) distribution. On the other hand, when the null hypothesis of interest is that the underlying process is an uncorrelated sequence, permutation tests are not necessarily level $α$, nor are they approximately level $α$ in large samples. In addition, permutation tests may have large Type 3, or directional, errors, in which a two-sided test rejects the null hypothesis and concludes that the first-order autocorrelation is larger than 0, when in fact it is less than 0. In this paper, under weak assumptions on the mixing coefficients and moments of the sequence, we provide a test procedure for which the asymptotic validity of the permutation test holds, while retaining the exact rejection probability $α$ in finite samples when the observations are independent and identically distributed. A Monte Carlo simulation study, comparing the permutation test to other tests of autocorrelation, is also performed, along with an empirical example of application to financial data.

math.ST

A New Approach for Large Scale Multiple Testing with Application to FDR Control for Graphically Structured Hypotheses

In many large scale multiple testing applications, the hypotheses often have a known graphical structure, such as gene ontology in gene expression data. Exploiting this graphical structure in multiple testing procedures can improve power as well as aid in interpretation. However, incorporating the structure into large scale testing procedures and proving that an error rate, such as the false discovery rate (FDR), is controlled can be challenging. In this paper, we introduce a new general approach for large scale multiple testing, which can aid in developing new procedures under various settings with proven control of desired error rates. This approach is particularly useful for developing FDR controlling procedures, which is simplified as the problem of developing per-family error rate (PFER) controlling procedures. Specifically, for testing hypotheses with a directed acyclic graph (DAG) structure, by using the general approach, under the assumption of independence, we first develop a specific PFER controlling procedure and based on this procedure, then develop a new FDR controlling procedure, which can preserve the desired DAG structure among the rejected hypotheses. Through a small simulation study and a real data analysis, we illustrate nice performance of the proposed FDR controlling procedure for DAG-structured hypotheses.

stat.ME

Control of Directional Errors in Fixed Sequence Multiple Testing

In this paper, we consider the problem of simultaneously testing many two-sided hypotheses when rejections of null hypotheses are accompanied by claims of the direction of the alternative. The fundamental goal is to construct methods that control the mixed directional familywise error rate, which is the probability of making any type 1 or type 3 (directional) error. In particular, attention is focused on cases where the hypotheses are ordered as $H_1 , \ldots, H_n$, so that $H_{i+1}$ is tested only if $H_1 , \ldots, H_i$ have all been previously rejected. In this situation, one can control the usual familywise error rate under arbitrary dependence by the basic procedure which tests each hypothesis at level $α$, and no other multiplicity adjustment is needed. However, we show that this is far too liberal if one also accounts for directional errors. But, by imposing certain dependence assumptions on the test statistics, one can retain the basic procedure. Through a simulation study and a clinical trial example, we numerically illustrate good performance of the proposed procedures compared to the existing mdFWER controlling procedures. The proposed procedures are also implemented in the R-package FixSeqMTP.

math.ST

Analysis of error control in large scale two-stage multiple hypothesis testing

When dealing with the problem of simultaneously testing a large number of null hypotheses, a natural testing strategy is to first reduce the number of tested hypotheses by some selection (screening or filtering) process, and then to simultaneously test the selected hypotheses. The main advantage of this strategy is to greatly reduce the severe effect of high dimensions. However, the first screening or selection stage must be properly accounted for in order to maintain some type of error control. In this paper, we will introduce a selection rule based on a selection statistic that is independent of the test statistic when the tested hypothesis is true. Combining this selection rule and the conventional Bonferroni procedure, we can develop a powerful and valid two-stage procedure. The introduced procedure has several nice properties: (i) it completely removes the selection effect; (ii) it reduces the multiplicity effect; (iii) it does not "waste" data while carrying out both selection and testing. Asymptotic power analysis and simulation studies illustrate that this proposed method can provide higher power compared to usual multiple testing methods while controlling the Type 1 error rate. Optimal selection thresholds are also derived based on our asymptotic analysis.

stat.ME

On Stepwise Control of Directional Errors under Independence and Some Dependence

In this paper, the problem of error control of stepwise multiple testing procedures is considered. For two-sided hypotheses, control of both type 1 and type 3 (or directional) errors is required, and thus mixed directional familywise error rate control and mixed directional false discovery rate control are each considered by incorporating both types of errors in the error rate. Mixed directional familywise error rate control of stepwise methods in multiple testing has proven to be a challenging problem, as demonstrated in Shaffer (1980). By an appropriate formulation of the problem, some new stepwise procedures are developed that control type 1 and directional errors under independence and various dependencies.

math.ST

Exact and asymptotically robust permutation tests

Given independent samples from P and Q, two-sample permutation tests allow one to construct exact level tests when the null hypothesis is P=Q. On the other hand, when comparing or testing particular parameters $θ$ of P and Q, such as their means or medians, permutation tests need not be level $α$, or even approximately level $α$ in large samples. Under very weak assumptions for comparing estimators, we provide a general test procedure whereby the asymptotic validity of the permutation test holds while retaining the exact rejection probability $α$ in finite samples when the underlying distributions are identical. The ideas are broadly applicable and special attention is given to the k-sample problem of comparing general parameters, whereby a permutation test is constructed which is exact level $α$ under the hypothesis of identical distributions, but has asymptotic rejection probability $α$ under the more general null hypothesis of equality of parameters. A Monte Carlo simulation study is performed as well. A quite general theory is possible based on a coupling construction, as well as a key contiguity argument for the multinomial and multivariate hypergeometric distributions.

math.ST

On the uniform asymptotic validity of subsampling and the bootstrap

This paper provides conditions under which subsampling and the bootstrap can be used to construct estimators of the quantiles of the distribution of a root that behave well uniformly over a large class of distributions $\mathbf{P}$. These results are then applied (i) to construct confidence regions that behave well uniformly over $\mathbf{P}$ in the sense that the coverage probability tends to at least the nominal level uniformly over $\mathbf{P}$ and (ii) to construct tests that behave well uniformly over $\mathbf{P}$ in the sense that the size tends to no greater than the nominal level uniformly over $\mathbf{P}$. Without these stronger notions of convergence, the asymptotic approximations to the coverage probability or size may be poor, even in very large samples. Specific applications include the multivariate mean, testing moment inequalities, multiple testing, the empirical process and U-statistics.

math.ST

Control of generalized error rates in multiple testing

Consider the problem of testing $s$ hypotheses simultaneously. The usual approach restricts attention to procedures that control the probability of even one false rejection, the familywise error rate (FWER). If $s$ is large, one might be willing to tolerate more than one false rejection, thereby increasing the ability of the procedure to correctly reject false null hypotheses. One possibility is to replace control of the FWER by control of the probability of $k$ or more false rejections, which is called the $k$-FWER. We derive both single-step and step-down procedures that control the $k$-FWER in finite samples or asymptotically, depending on the situation. We also consider the false discovery proportion (FDP) defined as the number of false rejections divided by the total number of rejections (and defined to be 0 if there are no rejections). The false discovery rate proposed by Benjamini and Hochberg [J. Roy. Statist. Soc. Ser. B 57 (1995) 289--300] controls $E(FDP)$. Here, the goal is to construct methods which satisfy, for a given $γ$ and $α$, $P\{FDP>γ\}\le α$, at least asymptotically. In contrast to the proposals of Lehmann and Romano [Ann. Statist. 33 (2005) 1138--1154], we construct methods that implicitly take into account the dependence structure of the individual test statistics in order to further increase the ability to detect false null hypotheses. This feature is also shared by related work of van der Laan, Dudoit and Pollard [Stat. Appl. Genet. Mol. Biol. 3 (2004) article 15], but our methodology is quite different. Like the work of Pollard and van der Laan [Proc. 2003 International Multi-Conference in Computer Science and Engineering, METMBS'03 Conference (2003) 3--9] and Dudoit, van der Laan and Pollard [Stat. Appl. Genet. Mol. Biol. 3 (2004) article 13], we employ resampling methods to achieve our goals. Some simulations compare finite sample performance to currently available methods.

math.ST

Stepup procedures for control of generalizations of the familywise error rate

Consider the multiple testing problem of testing null hypotheses $H_1,...,H_s$. A classical approach to dealing with the multiplicity problem is to restrict attention to procedures that control the familywise error rate ($\mathit{FWER}$), the probability of even one false rejection. But if $s$ is large, control of the $\mathit{FWER}$ is so stringent that the ability of a procedure that controls the $\mathit{FWER}$ to detect false null hypotheses is limited. It is therefore desirable to consider other measures of error control. This article considers two generalizations of the $\mathit{FWER}$. The first is the $k-\mathit{FWER}$, in which one is willing to tolerate $k$ or more false rejections for some fixed $k\geq 1$. The second is based on the false discovery proportion ($\mathit{FDP}$), defined to be the number of false rejections divided by the total number of rejections (and defined to be 0 if there are no rejections). Benjamini and Hochberg [J. Roy. Statist. Soc. Ser. B 57 (1995) 289--300] proposed control of the false discovery rate ($\mathit{FDR}$), by which they meant that, for fixed $α$, $E(\mathit{FDP})\leqα$. Here, we consider control of the $\mathit{FDP}$ in the sense that, for fixed $γ$ and $α$, $P\{\mathit{FDP}>γ\}\leq α$. Beginning with any nondecreasing sequence of constants and $p$-values for the individual tests, we derive stepup procedures that control each of these two measures of error control without imposing any assumptions on the dependence structure of the $p$-values. We use our results to point out a few interesting connections with some closely related stepdown procedures. We then compare and contrast two $\mathit{FDP}$-controlling procedures obtained using our results with the stepup procedure for control of the $\mathit{FDR}$ of Benjamini and Yekutieli [Ann. Statist. 29 (2001) 1165--1188].

math.ST