SearcharxivSearch

arXiv subjects

Mengsi Gao

Publications and source records attributed to Mengsi Gao.

7 recordsLinked to original sources

Misspecified regressions with mixed regressors: robust inference and causal interpretation

For analytic convenience, existing statistical frameworks either assume random or fixed regressors. However, it is a little awkward that they do not cover the practical case of estimating the average treatment effect in experiments with randomized treatments and non-randomized, fixed pretreatment covariates. We unify the literature by providing the theory for regressions with mixed regressors that contain both random and fixed components. Importantly, our theory allows for misspecification of the regression functions. We first establish general results for estimating equations with both random and fixed components and then use it to analyze misspecified linear regression, with applications to completely randomized experiments. We focus on the causal interpretation of the regression coefficients and standard errors even when the models are wrong. We start with the theory for independent data and then extend the discussion to clustered data.

math.ST

Coupling and Maximal Inequalities for Graph-Dependent Empirical Processes

We develop maximal inequalities for empirical processes indexed by graph-dependent observations. Our bounds separate the complexity of the indexing class from two features specific to graph dependence: the geometry of the underlying graph and the cost of coupling graph-separated blocks to independent copies. The coupling construction combines a novel graph-adapted dependence coefficient with a coloring of a block partition. We specialize the results to graphs with polynomial and exponential growth and to directed dyadic graphs. We then derive Glivenko--Cantelli results and characterize the associated effective sample size. A central implication is that graph-dependent empirical processes need not exhibit a generic root-$n$ rate: convergence is jointly determined by function-class complexity, graph geometry, and the decay of dependence with graph distance. Finally, we apply the results to obtain uniform laws of large numbers for network autoregressive models, nonlinear local-propagation models, and treatment-interference settings.

math.PR

Causal inference in network experiments: regression-based analysis and design-based properties

Network experiments are powerful tools for studying spillover effects, which avoid endogeneity by randomly assigning treatments to units over networks. However, it is non-trivial to analyze network experiments properly without imposing strong modeling assumptions. We show that regression-based point estimators and standard errors can have strong theoretical guarantees if the regression functions and robust standard errors are carefully specified to accommodate the interference patterns under network experiments. We first recall a well-known result that the Hájek estimator is numerically identical to the coefficient from the weighted-least-squares fit based on the inverse probability of the exposure mapping. Moreover, we demonstrate that the regression-based approach offers three notable advantages: its ease of implementation, the ability to derive standard errors through the same regression fit, and the potential to integrate covariates into the analysis to improve efficiency. Recognizing that the regression-based network-robust covariance estimator can be anti-conservative under nonconstant effects, we propose an adjusted covariance estimator to improve the empirical coverage rates.

econ.EM

Identification and Inference on Treatment Effects under Covariate-Adaptive Randomization and Imperfect Compliance

Randomized controlled trials (RCTs) frequently utilize covariate-adaptive randomization (CAR) (e.g., stratified block randomization) and commonly suffer from imperfect compliance. This paper studies the identification and inference for the average treatment effect (ATE) and the average treatment effect on the treated (ATT) in such RCTs with a binary treatment. We first develop characterizations of the identified sets for both estimands. Since data are generally not i.i.d. under CAR, these characterizations do not follow from existing results. We then provide consistent estimators of the identified sets and asymptotically valid confidence intervals for the parameters. Our asymptotic analysis leads to concrete practical recommendations regarding how to estimate the treatment assignment probabilities that enter the estimated bounds. For the ATE bounds, using sample analog assignment frequencies is more efficient than relying on the true assignment probabilities. For the ATT bounds, the most efficient approach is to use the true assignment probability for the probabilities in the numerator and the sample analog for those in the denominator.

econ.EM

Endogenous Interference in Randomized Experiments

This paper investigates the identification and inference of treatment effects in randomized controlled trials with social interactions. Two key network features characterize the setting and introduce endogeneity: (1) latent variables may affect both network formation and outcomes, and (2) the intervention may alter network structure, mediating treatment effects. I make three contributions. First, I define parameters within a post-treatment network framework, distinguishing direct effects of treatment from indirect effects mediated through changes in network structure. I provide a causal interpretation of the coefficients in a linear outcome model. For estimation and inference, I focus on a specific form of peer effects, represented by the fraction of treated friends. Second, in the absence of endogeneity, I establish the consistency and asymptotic normality of ordinary least squares estimators. Third, if endogeneity is present, I propose addressing it through shift-share instrumental variables, demonstrating the consistency and asymptotic normality of instrumental variable estimators in relatively sparse networks. For denser networks, I propose a denoised estimator based on eigendecomposition to restore consistency. Finally, I revisit Prina (2015) as an empirical illustration, demonstrating that treatment can influence outcomes both directly and through network structure changes.

econ.EM

On the power properties of inference for parameters with interval identified sets

This paper studies the power properties of confidence intervals (CIs) for a partially-identified parameter of interest with an interval identified set. We assume the researcher has bounds estimators needed to construct the CIs proposed by Imbens and Manski (2004), Stoye (2009), and Stoye (2020), denoted by CI_alpha^1, CI_alpha^2, CI_alpha^3, and CI_alpha^4. We also assume these bounds estimators are ``ordered'': the lower bound estimator is less than or equal to the upper bound estimator. This setup arises in economic applications involving missing data and treatment effects. Under these conditions, we establish two results. First, we show that CI_alpha^1 and CI_alpha^2 are equally powerful, and both dominate CI_alpha^3 and CI_alpha^4. Second, we consider a favorable situation in which there are two possible bounds estimators to construct these CIs, and one is more efficient than the other. One would expect that the more efficient bounds estimator yields more powerful inference. We prove that this desirable result holds for CI_alpha^1 and CI_alpha^2, but not necessarily for CI_alpha^3 or CI_alpha^4. In summary, within the class of models considered, CI_alpha^1 and CI_alpha^2 have identical power properties, and both compare favorably to CI_alpha^3 or CI_alpha^4.

econ.EM

Inference under Covariate-Adaptive Randomization with Imperfect Compliance

This paper studies inference in a randomized controlled trial (RCT) with covariate-adaptive randomization (CAR) and imperfect compliance of a binary treatment. In this context, we study inference on the LATE. As in Bugni et al. (2018,2019), CAR refers to randomization schemes that first stratify according to baseline covariates and then assign treatment status so as to achieve "balance" within each stratum. In contrast to these papers, however, we allow participants of the RCT to endogenously decide to comply or not with the assigned treatment status. We study the properties of an estimator of the LATE derived from a "fully saturated" IV linear regression, i.e., a linear regression of the outcome on all indicators for all strata and their interaction with the treatment decision, with the latter instrumented with the treatment assignment. We show that the proposed LATE estimator is asymptotically normal, and we characterize its asymptotic variance in terms of primitives of the problem. We provide consistent estimators of the standard errors and asymptotically exact hypothesis tests. In the special case when the target proportion of units assigned to each treatment does not vary across strata, we can also consider two other estimators of the LATE, including the one based on the "strata fixed effects" IV linear regression, i.e., a linear regression of the outcome on indicators for all strata and the treatment decision, with the latter instrumented with the treatment assignment. Our characterization of the asymptotic variance of the LATE estimators allows us to understand the influence of the parameters of the RCT. We use this to propose strategies to minimize their asymptotic variance in a hypothetical RCT based on data from a pilot study. We illustrate the practical relevance of these results using a simulation study and an empirical application based on Dupas et al. (2018).

econ.EM