Searcharxiv⌕ Search

arXiv subjects

Azeem M. Shaikh

Publications and source records attributed to Azeem M. Shaikh.

At least 19 recordsLinked to original sources

Inference for Treatment Effects Conditional on Generalized Principal Strata using Instrumental Variables

We propose a general approach to inference for a broad class of models that arise in the analysis of treatment effects with discrete-valued treatments and instruments and a general-valued outcome. In addition to instrument exogeneity, the main substantive assumption in our class of models rules out certain response types by assuming that they occur with probability zero. Here, the response type refers to the vector of potential outcomes and potential treatments, and we refer to a set of possible values for the response type as a generalized principal stratum. Through a series of examples, we show that this framework encompasses a wide variety of assumptions that have been considered in the previous literature. Our framework allows inference on any treatment effect parameter that can be expressed as the expectation of a function of the response type conditional on a generalized principal stratum. We develop methods for inference on such parameters under these assumptions, as well as methods for testing the validity of the assumptions themselves. A key result of our analysis is a characterization of the identified set for such parameters under these assumptions and the testable restrictions for the assumptions themselves in terms of existence of a nonnegative solution to linear systems of equations with a special structure. We propose methods for inference exploiting this special structure and recent results in Fang et al. (2023).

econ.EM↗

Graph-Laplacian Variance Estimators for Finely Stratified Experiments

This paper considers design-based inference on the average treatment effect in finely stratified experiments, where uncertainty arises only from the randomized treatment assignment. We focus on settings in which units are first stratified into groups of fixed size according to baseline covariates and, then within each group, exactly one unit is assigned to treatment. In this setting, we introduce a class of graph-Laplacian variance estimators in which strata form the vertices of a weighted graph and edge weights determine how between-stratum comparisons are aggregated. The canonical estimator of Imai (2008) corresponds to a complete graph with edge weights normalized so that each stratum has weighted degree one, while a paired-stratum estimator arises from a perfect matching graph. For the subclass of degree-calibrated graphs, in which each vertex has weighted degree one, we derive an exact bias identity showing that the corresponding estimators are upward-biased, with bias governed by squared differences in the true stratum-level treatment effects across adjacent strata. As a result, any such estimator may be used for valid inference. The identity further suggests that paired-stratum estimators constructed from a covariate-based perfect matching can induce small biases when treatment effects vary smoothly with the covariates. Without such smoothness, however, we show that paired-stratum estimators can exhibit large worst-case bias, and that, within the class of degree-calibrated estimators, the complete-graph estimator is minimax optimal for normalized bias under a weak bound on treatment-effect heterogeneity. Motivated by this contrast, we propose a regularized graph estimator that controls worst-case normalized bias while preserving much of the locality of the paired-stratum estimator. Simulations illustrate the resulting tradeoff between locality and worst-case protection.

econ.EM↗

Randomization Tests in Randomized Saturation Designs

Randomized saturation designs are widely used to study spillover effects in clustered populations. In these designs, clusters are first assigned to treatment saturation levels, and units are then randomized within clusters according to the assigned saturation. This paper develops randomization tests for such experiments under several null hypotheses that arise naturally in spillover analysis. For a fixed pair of saturation levels, we first study two individual-level hypotheses: a partially sharp null of no spillover effect for every untreated unit and a bounded null that restricts individual spillover effects by a prespecified constant. Both hypotheses can be tested using a common conditional randomization framework, with finite-sample validity obtained by combining the same focal-unit relabeling distribution with null-specific statistics. We then study weak average-spillover nulls and show that, although these nulls do not yield finite-sample exact conditional tests, studentized relabeling statistics deliver asymptotically valid randomization-based inference. Finally, for multiple ordered saturation levels, we develop a finite-sample valid unconditional pairwise-imputation test for global monotonicity of spillover effects. Simulations and an application to the Zomba Cash Transfer experiment illustrate the finite-sample behavior and practical implementation of the methods.

stat.ME↗

Inference for Linear Systems with Unknown Coefficients

This paper considers the problem of testing whether there exists a solution satisfying certain non-negativity constraints to a linear system of equations. Importantly and in contrast to some prior work, we allow all parameters in the system of equations, including the slope coefficients, to be unknown. For this reason, we describe the linear system as having unknown (as opposed to known) coefficients. This hypothesis testing problem arises naturally when constructing confidence sets for possibly partially identified parameters in the analysis of nonparametric instrumental variables models, treatment effect models, and random coefficient models, among other settings. To rule out certain instances in which the testing problem is impossible, in the sense that the power of any test will be bounded by its size, we begin our analysis by characterizing the closure of the null hypothesis with respect to the total variation distance. We then use this characterization to develop novel testing procedures based on sample-splitting. We establish the validity of our testing procedures under weak and interpretable conditions on the linear system. An important feature of these conditions is that they permit the dimensionality of the problem to grow rapidly with the sample size. A further attractive property of our tests is that they do not require simulation to compute suitable critical values. We illustrate the practical relevance of our theoretical results in a simulation study.

econ.EM↗

Randomization Inference in Two-Sided Market Experiments

Randomized experiments are increasingly employed in two-sided markets, such as buyer--seller platforms, to evaluate the effects of marketplace interventions. These experiments must reflect the underlying two-sided market structure in their design and can therefore be challenging to analyze. In this paper, we develop a randomization inference framework for outcomes from two-sided experiments, with a focus on testing and inference for two-sided spillover effects. Our approach is finite-sample valid under sharp null hypotheses. Regarding weak null hypotheses, we find that the commonly used Neyman-style studentization does not universally ensure asymptotic validity, and we document how it depends on the specific formulation of the null. We then propose a two-way variance estimator for studentization that restores asymptotic validity. We further propose methods to improve testing power by exploiting the two-sided structure of the problem, which we validate empirically. We demonstrate our methods through a series of simulation studies and an applied example from a network experiment in micro-lending.

stat.ME↗

On the Efficiency of Highly Stratified Experiments

This paper studies the use of highly stratified designs for the efficient estimation of a large class of treatment effect parameters that arise in the analysis of experiments. By a "highly stratified" design, we mean experiments in which units are divided into blocks of a fixed size and a proportion within each block is assigned to a binary treatment uniformly at random. The class of parameters considered are those that can be expressed as the solution to a set of moment conditions constructed using a known function of the observed data. They include, among other things, average treatment effects, quantile treatment effects, and local average treatment effects as well as the counterparts to these quantities in experiments in which the unit is itself a cluster. In this setting, we establish three results. First, we show that under a highly stratified design, the naïve method of moments estimator achieves the same asymptotic variance as what could typically be attained under alternative treatment assignment mechanisms only through ex post covariate adjustment. Second, we argue that the naïve method of moments estimator under a highly stratified design is asymptotically efficient by deriving a lower bound on the asymptotic variance of regular estimators of the parameter of interest in the form of a convolution theorem. In this sense, highly stratified experiments are attractive because they lead to efficient estimators of treatment effect parameters "by design." Finally, we strengthen this conclusion by establishing conditions under which a "fast-balancing" property of highly stratified designs is in fact necessary for the naïve method of moments estimator to attain the efficiency bound.

econ.EM↗

Inference After Ranking with Applications to Economic Mobility

This paper considers the problem of inference after ranking. In our setting, we are interested in any population whose rank according to some random quantity, such as an estimated treatment effect, a measure of value-added, or benefit (net of cost), falls in a pre-specified range of values. As such, this framework generalizes the inference on winners setting previously considered in Andrews et al. (2023), in which a winner is understood to be the single population whose rank according to some random quantity is highest. We show that this richer setting accommodates a broad variety of empirically-relevant applications. We develop a two-step method for inference, which we compare to existing methods or their natural generalizations to this setting. We first show the finite-sample validity of this method in a normal location model and then develop asymptotic counterparts to these results by proving uniform validity over a large class of distributions satisfying a weak uniform integrability condition. Importantly, our results permit degeneracy in the covariance matrix of the limiting distribution, which arises naturally in many applications. In an application to the literature on economic mobility, we find that it is difficult to distinguish between high and low-mobility census tracts when correcting for selection. Finally, we demonstrate the practical relevance of our theoretical results through an extensive set of simulations.

econ.EM↗

On the Identifying Power of Generalized Monotonicity for Average Treatment Effects

In the context of a binary outcome, treatment, and instrument, Balke and Pearl (1993, 1997) es- tablish that the monotonicity condition of Imbens and Angrist (1994) has no identifying power beyond instrument exogeneity for average potential outcomes and average treatment effects in the sense that adding it to instrument exogeneity does not decrease the identified sets for those parameters whenever those restrictions are consistent with the distribution of the observable data. This paper shows that this phenomenon holds in a broader setting with a multi-valued outcome, treatment, and instrument, under an extension of the monotonicity condition that we refer to as generalized monotonicity. We further show that this phenomenon holds for any restriction on treatment response that is stronger than generalized monotonicity provided that these stronger restrictions do not restrict potential outcomes. Importantly, many models of potential treatments previously considered in the literature imply generalized monotonic- ity, including the types of monotonicity restrictions considered by Kline and Walters (2016), Kirkeboen et al. (2016), and Heckman and Pinto (2018), and the restriction that treatment selection is determined by particular classes of additive random utility models. We show through a series of examples that restrictions on potential treatments can provide identifying power beyond instrument exogeneity for av- erage potential outcomes and average treatment effects when the restrictions imply that the generalized monotonicity condition is violated. In this way, our results shed light on the types of restrictions required for help in identifying average potential outcomes and average treatment effects.

econ.EM↗

Reasonable uncertainty: Confidence intervals in empirical Bayes discrimination detection

We revisit empirical Bayes discrimination detection, focusing on uncertainty arising from both partial identification and sampling variability. While prior work has mostly focused on partial identification, we find that some empirical findings are not robust to sampling uncertainty. To better connect statistical evidence to the magnitude of real-world discriminatory behavior, we propose a counterfactual odds-ratio estimand with a attractive properties and interpretation. Our analysis reveals the importance of careful attention to uncertainty quantification and downstream goals in empirical Bayes analyses.

econ.EM↗

Inference in Cluster Randomized Trials with Matched Pairs

This paper studies inference in cluster randomized trials where treatment status is determined according to a "matched pairs" design. Here, by a cluster randomized experiment, we mean one in which treatment is assigned at the level of the cluster; by a "matched pairs" design, we mean that a sample of clusters is paired according to baseline, cluster-level covariates and, within each pair, one cluster is selected at random for treatment. We study the large-sample behavior of a weighted difference-in-means estimator and derive two distinct sets of results depending on if the matching procedure does or does not match on cluster size. We then propose a single variance estimator which is consistent in either regime. Combining these results establishes the asymptotic exactness of tests based on these estimators. Next, we consider the properties of two common testing procedures based on t-tests constructed from linear regressions, and argue that both are generally conservative in our framework. We additionally study the behavior of a randomization test which permutes the treatment status for clusters within pairs, and establish its finite-sample and asymptotic validity for testing specific null hypotheses. Finally, we propose a covariate-adjusted estimator which adjusts for additional baseline covariates not used for treatment assignment, and establish conditions under which such an estimator leads to strict improvements in precision. A simulation study confirms the practical relevance of our theoretical results.

econ.EM↗

A New Design-Based Variance Estimator for Finely Stratified Experiments

This paper considers the problem of design-based inference for the average treatment effect in finely stratified experiments. Here, by "design-based'' we mean that the only source of uncertainty stems from the randomness in treatment assignment; by "finely stratified'' we mean that units are stratified into groups of a fixed size according to baseline covariates and then, within each group, a fixed number of units are assigned uniformly at random to treatment and the remainder to control. In this setting we present a novel estimator of the variance of the difference-in-means based on pairing "adjacent" strata. Importantly, our estimator is well defined even in the challenging setting where there is exactly one treated or control unit per stratum. We prove that our estimator is upward-biased, and thus can be used for inference under mild restrictions on the finite population. We compare our estimator with some well-known estimators that have been proposed previously in this setting, and demonstrate that, while these estimators are also upward-biased, our estimator has smaller bias and therefore leads to more precise inferences whenever adjacent strata are sufficiently similar. To further understand when our estimator leads to more precise inferences, we introduce a framework motivated by a thought experiment in which the finite population is modeled as having been drawn once in an i.i.d. fashion from a well-behaved probability distribution. In this framework, we argue that our estimator dominates the others in terms of limiting bias and that these improvements are strict except under strong restrictions on the treatment effects. Finally, we illustrate the practical relevance of our theoretical results through a simulation study, which reveals that our estimator can in fact lead to substantially more precise inferences, especially when the quality of stratification is high.

econ.EM↗

A Primer on the Analysis of Randomized Experiments and a Survey of some Recent Advances

The past two decades have witnessed a surge of new research in the analysis of randomized experiments. The emergence of this literature may seem surprising given the widespread use and long history of experiments as the "gold standard" in program evaluation, but this body of work has revealed many subtle aspects of randomized experiments that may have been previously unappreciated. This article provides an overview of some of these topics, primarily focused on stratification, regression adjustment, and cluster randomization.

econ.EM↗

Randomization Inference: Theory and Applications

We review approaches to statistical inference based on randomization. Permutation tests are treated as an important special case. Under a certain group invariance property, referred to as the ``randomization hypothesis,'' randomization tests achieve exact control of the Type I error rate in finite samples. Although this unequivocal precision is very appealing, the range of problems that satisfy the randomization hypothesis is somewhat limited. We show that randomization tests are often asymptotically, or approximately, valid and efficient in settings that deviate from the conditions required for finite-sample error control. When randomization tests fail to offer even asymptotic Type 1 error control, their asymptotic validity may be restored by constructing an asymptotically pivotal test statistic. Randomization tests can then provide exact error control for tests of highly structured hypotheses with good performance in a wider class of problems. We give a detailed overview of several prominent applications of randomization tests, including two-sample permutation tests, regression, and conformal inference.

econ.EM↗

Inference in Experiments with Matched Pairs and Imperfect Compliance

This paper studies inference for the local average treatment effect in randomized controlled trials with imperfect compliance where treatment status is determined according to "matched pairs." By "matched pairs," we mean that units are sampled i.i.d. from the population of interest, paired according to observed, baseline covariates and finally, within each pair, one unit is selected at random for treatment. Under weak assumptions governing the quality of the pairings, we first derive the limit distribution of the usual Wald (i.e., two-stage least squares) estimator of the local average treatment effect. We show further that conventional heteroskedasticity-robust estimators of the Wald estimator's limiting variance are generally conservative, in that their probability limits are (typically strictly) larger than the limiting variance. We therefore provide an alternative estimator of the limiting variance that is consistent. Finally, we consider the use of additional observed, baseline covariates not used in pairing units to increase the precision with which we can estimate the local average treatment effect. To this end, we derive the limiting behavior of a two-stage least squares estimator of the local average treatment effect which includes both the additional covariates in addition to pair fixed effects, and show that its limiting variance is always less than or equal to that of the Wald estimator. To complete our analysis, we provide a consistent estimator of this limiting variance. A simulation study confirms the practical relevance of our theoretical results. Finally, we apply our results to revisit a prominent experiment studying the effect of macroinsurance on microenterprise in Egypt.

econ.EM↗

Covariate Adjustment in Experiments with Matched Pairs

This paper studies inference on the average treatment effect in experiments in which treatment status is determined according to "matched pairs" and it is additionally desired to adjust for observed, baseline covariates to gain further precision. By a "matched pairs" design, we mean that units are sampled i.i.d. from the population of interest, paired according to observed, baseline covariates and finally, within each pair, one unit is selected at random for treatment. Importantly, we presume that not all observed, baseline covariates are used in determining treatment assignment. We study a broad class of estimators based on a "doubly robust" moment condition that permits us to study estimators with both finite-dimensional and high-dimensional forms of covariate adjustment. We find that estimators with finite-dimensional, linear adjustments need not lead to improvements in precision relative to the unadjusted difference-in-means estimator. This phenomenon persists even if the adjustments are interacted with treatment; in fact, doing so leads to no changes in precision. However, gains in precision can be ensured by including fixed effects for each of the pairs. Indeed, we show that this adjustment is the "optimal" finite-dimensional, linear adjustment. We additionally study two estimators with high-dimensional forms of covariate adjustment based on the LASSO. For each such estimator, we show that it leads to improvements in precision relative to the unadjusted difference-in-means estimator and also provide conditions under which it leads to the "optimal" nonparametric, covariate adjustment. A simulation study confirms the practical relevance of our theoretical analysis, and the methods are employed to reanalyze data from an experiment using a "matched pairs" design to study the effect of macroinsurance on microenterprise.

econ.EM↗

On the implementation of Approximate Randomization Tests in Linear Models with a Small Number of Clusters

This paper provides a user's guide to the general theory of approximate randomization tests developed in Canay, Romano, and Shaikh (2017) when specialized to linear regressions with clustered data. An important feature of the methodology is that it applies to settings in which the number of clusters is small -- even as small as five. We provide a step-by-step algorithmic description of how to implement the test and construct confidence intervals for the parameter of interest. In doing so, we additionally present three novel results concerning the methodology: we show that the method admits an equivalent implementation based on weighted scores; we show the test and confidence intervals are invariant to whether the test statistic is studentized or not; and we prove convexity of the confidence intervals for scalar parameters. We also articulate the main requirements underlying the test, emphasizing in particular common pitfalls that researchers may encounter. Finally, we illustrate the use of the methodology with two applications that further illuminate these points. The companion {\tt R} and {\tt Stata} packages facilitate the implementation of the methodology and the replication of the empirical exercises.

econ.EM↗

Confidence Intervals for Seroprevalence

This paper concerns the construction of confidence intervals in standard seroprevalence surveys. In particular, we discuss methods for constructing confidence intervals for the proportion of individuals in a population infected with a disease using a sample of antibody test results and measurements of the test's false positive and false negative rates. We begin by documenting erratic behavior in the coverage probabilities of standard Wald and percentile bootstrap intervals when applied to this problem. We then consider two alternative sets of intervals constructed with test inversion. The first set of intervals are approximate, using either asymptotic or bootstrap approximation to the finite-sample distribution of a chosen test statistic. We consider several choices of test statistic, including maximum likelihood estimators and generalized likelihood ratio statistics. We show with simulation that, at empirically relevant parameter values and sample sizes, the coverage probabilities for these intervals are close to their nominal level and are approximately equi-tailed. The second set of intervals are shown to contain the true parameter value with probability at least equal to the nominal level, but can be conservative in finite samples.

stat.AP↗

Inference for Large-Scale Linear Systems with Known Coefficients

This paper considers the problem of testing whether there exists a non-negative solution to a possibly under-determined system of linear equations with known coefficients. This hypothesis testing problem arises naturally in a number of settings, including random coefficient, treatment effect, and discrete choice models, as well as a class of linear programming problems. As a first contribution, we obtain a novel geometric characterization of the null hypothesis in terms of identified parameters satisfying an infinite set of inequality restrictions. Using this characterization, we devise a test that requires solving only linear programs for its implementation, and thus remains computationally feasible in the high-dimensional applications that motivate our analysis. The asymptotic size of the proposed test is shown to equal at most the nominal level uniformly over a large class of distributions that permits the number of linear equations to grow with the sample size.

econ.EM↗