Searcharxiv⌕ Search

arXiv subjects

Peter Z. Schochet

Publications and source records attributed to Peter Z. Schochet.

5 recordsLinked to original sources

Design-Based RCT Estimators and Central Limit Theorems for Baseline Subgroup and Related Analyses

There is a growing literature on design-based methods to estimate average treatment effects (ATEs) for randomized controlled trials (RCTs) for full sample analyses. This article extends these methods to estimate ATEs for discrete subgroups defined by pre-treatment variables, with an application to an RCT testing subgroup effects for a school voucher experiment in New York City. We consider ratio estimators for subgroup effects using regression methods, allowing for model covariates to improve precision, and prove a finite population central limit theorem. We discuss extensions to blocked and clustered RCT designs, and to other common estimators with random treatment-control sample sizes (or weights): post-stratification estimators, weighted estimators that adjust for data nonresponse, and estimators for Bernoulli trials. We also develop simple variance estimators that share features with robust estimators. Simulations show that the design-based subgroup estimators yield confidence interval coverage near nominal levels, even for small subgroups.

stat.ME↗

Estimating Complier Average Causal Effects for Clustered RCTs When the Treatment Affects the Service Population

RCTs sometimes test interventions that aim to improve existing services targeted to a subset of individuals identified after randomization. Accordingly, the treatment could affect the composition of service recipients and the offered services. With such bias, intention-to-treat estimates using data on service recipients and nonrecipients may be difficult to interpret. This article develops causal estimands and inverse probability weighting (IPW) estimators for complier populations in these settings, using a generalized estimating equation approach that adjusts the standard errors for estimation error in the IPW weights. While our focus is on more general clustered RCTs, the methods also apply (reduce) to non-clustered RCTs. Simulations show that the estimators achieve nominal confidence interval coverage under the assumed identification conditions. An empirical application demonstrates the methods using data from a large-scale RCT testing the effects of early childhood services on children's cognitive development scores.

stat.ME↗

Statistical Power for Estimating Treatment Effects Using Difference-in-Differences and Comparative Interrupted Time Series Designs with Variation in Treatment Timing

This article develops new closed-form variance expressions for power analyses for commonly used difference-in-differences (DID) and comparative interrupted time series (CITS) panel data estimators. The main contribution is to incorporate variation in treatment timing into the analysis. The power formulas also account for other key design features that arise in practice: autocorrelated errors, unequal measurement intervals, and clustering due to the unit of treatment assignment. We consider power formulas for both cross-sectional and longitudinal models and allow for covariates. An illustrative power analysis provides guidance on appropriate sample sizes. The key finding is that accounting for treatment timing increases required sample sizes. Further, DID estimators have considerably more power than standard CITS and ITS estimators. An available Shiny R dashboard performs the sample size calculations for the considered estimators.

stat.ME↗

Design-Based Ratio Estimators and Central Limit Theorems for Clustered, Blocked RCTs

This article develops design-based ratio estimators for clustered, blocked randomized controlled trials (RCTs), with an application to a federally funded, school-based RCT testing the effects of behavioral health interventions. We consider finite population weighted least squares estimators for average treatment effects (ATEs), allowing for general weighting schemes and covariates. We consider models with block-by-treatment status interactions as well as restricted models with block indicators only. We prove new finite population central limit theorems for each block specification. We also discuss simple variance estimators that share features with commonly used cluster-robust standard error estimators. Simulations show that the design-based ATE estimator yields nominal rejection rates with standard errors near true ones, even with few clusters.

stat.ME↗

A Lasso-OLS Hybrid Approach to Covariate Selection and Average Treatment Effect Estimation for Clustered RCTs Using Design-Based Methods

Statistical power is often a concern for clustered RCTs due to variance inflation from design effects and the high cost of adding study clusters (such as hospitals, schools, or communities). While covariate pre-specification is the preferred approach for improving power to estimate regression-adjusted average treatment effects (ATEs), further precision gains can be achieved through covariate selection once primary outcomes have been collected. This article uses design-based methods underlying clustered RCTs to develop a Lasso-OLS hybrid procedure for the post-hoc selection of covariates and ATE estimation that avoids model overfitting and lack of transparency. In the first stage, lasso estimation is conducted using cluster-level averages, where asymptotic normality is proved using a new central limit theorem for finite population regression estimators. In the second stage, ATEs and design-based standard errors are estimated using weighted least squares with the first stage lasso covariates. This nonparametric approach applies to continuous, binary, and discrete outcomes. Simulation results indicate that Type 1 errors of the second stage ATE estimates are near nominal values and standard errors are near true ones, although somewhat conservative with small samples. The method is demonstrated using data from a large, federally funded clustered RCT testing the effects of school-based programs promoting behavioral health.

stat.ME↗