SearcharxivSearch

arXiv subjects

Max Cytrynbaum

Publications and source records attributed to Max Cytrynbaum.

6 recordsLinked to original sources

The Limits of Experimental Design: Covariate Balance Beyond Low Dimension

We study how fast experimental designs can approach the semiparametric efficiency bound in finite samples, as measured by the excess variance of unadjusted treatment effect estimation. We prove an impossibility theorem: under weak conditions, no design can approach the variance bound uniformly over smooth outcome models unless covariate dimension $d \ll \log n$. Even in experiments with thousands of units, this permits only a handful of covariates. Motivated by this, we propose new designs based on discrepancy minimization that instead attempt to control imbalances over restricted-complexity nonparametric function classes. Such designs achieve fast rates to their corresponding restricted efficiency targets, permitting $d \ll n$ covariates in an additive nonparametric specification. They can also be combined with matching to protect against unmodeled outcome variation. In simulations calibrated to 12 published experiments, our designs reduce variance relative to matched pairs randomization in every empirical setting.

econ.EM

Coupling Designs for Randomized Experiments with Complex Treatments

We describe a new family of experimental designs that extends the principle of stratified randomization to settings with continuous, constrained multivariate, and other irregular treatment spaces. Our approach is to first match units into homogeneous groups, then use Monte Carlo couplings to assign within-group treatments to be highly dispersed over the treatment space. We show that ensuring similar units receive dissimilar treatments improves estimation efficiency. The efficiency gains are proportional to the product of dispersion and match quality, where dispersion measures how spread out the assignments are relative to independent randomization. We develop a new spectral analysis showing how efficiency depends on alignment between the smoothness and shape of the estimator's influence function and the coupling's principal directions. We illustrate these designs with examples from development, behavioral, and labor economics. In particular, our empirical application uses data from a real experiment allocating savings monitors using their position within village social networks.

econ.EM

Finely Stratified Rerandomization Designs

Finely stratified randomization makes unadjusted treatment effect estimation efficient when matches are tight. However, match quality deteriorates rapidly as covariate dimension increases, attenuating the gains from stratification. Motivated by this problem, we study designs that finely stratify on a few important covariates, then rerandomize within matched groups to balance the remainder. We derive the asymptotic distribution for GMM estimators of general causal parameters under such designs, showing that they provide nonparametric control over the stratification covariates and linear control over the rerandomization covariates. The resulting distribution is generally non-normal, but optimal linear adjustment restores asymptotic normality. For finite population parameters, we derive upper bounds on the non-identified asymptotic variance, enabling conservative inference that accounts for the efficiency gains from both design stages. An empirical application to estimating treatment effect heterogeneity among compliers illustrates the gains from adding rerandomization to a stratified design.

econ.EM

Covariate Adjustment in Stratified Experiments

This paper studies covariate adjusted estimation of the average treatment effect in stratified experiments. We work in a general framework that includes matched tuples designs, coarse stratification, and complete randomization as special cases. Regression adjustment with treatment-covariate interactions is known to weakly improve efficiency for completely randomized designs. By contrast, we show that for stratified designs such regression estimators are generically inefficient, potentially even increasing estimator variance relative to the unadjusted benchmark. Motivated by this result, we derive the asymptotically optimal linear covariate adjustment for a given stratification. We construct several feasible estimators that implement this efficient adjustment in large samples. In the special case of matched pairs, for example, the regression including treatment, covariates, and pair fixed effects is asymptotically optimal. We also provide novel asymptotically exact inference methods that allow researchers to report smaller confidence intervals, fully reflecting the efficiency gains from both stratification and adjustment. Simulations and an empirical application demonstrate the value of our proposed methods.

econ.EM

Fine Stratification of Survey Experiments

This paper studies a two-stage model of experimentation, where the researcher first samples representative experimental participants from an eligible pool, then assigns each sampled unit to treatment or control, using matched $k$-tuples randomization at both stages. To implement such designs, we develop a fast new algorithm for matching units into $k$-tuples for any $k \ge 2$ and any dimension of covariates. By surveying 200 recent experimental working papers, we estimate that our algorithm newly enables multivariate fine stratification with provable match quality guarantees for about 44\% of experiments in economics. We show that finely stratified sampling and assignment both nonparametrically reduce the variance of treatment effect estimation, with the gains from stratified sampling increasing in the size of the eligible pool and how well covariates predict treatment effect heterogeneity. We develop new inference methods that fully exploit the efficiency gains from both design stages, allowing researchers to report smaller standard errors if they designed a representative experiment. An application to nine published experiments quantifies the efficiency gains.

econ.EM

Blocked Clusterwise Regression

A recent literature in econometrics models unobserved cross-sectional heterogeneity in panel data by assigning each cross-sectional unit a one-dimensional, discrete latent type. Such models have been shown to allow estimation and inference by regression clustering methods. This paper is motivated by the finding that the clustered heterogeneity models studied in this literature can be badly misspecified, even when the panel has significant discrete cross-sectional structure. To address this issue, we generalize previous approaches to discrete unobserved heterogeneity by allowing each unit to have multiple, imperfectly-correlated latent variables that describe its response-type to different covariates. We give inference results for a k-means style estimator of our model and develop information criteria to jointly select the number clusters for each latent variable. Monte Carlo simulations confirm our theoretical results and give intuition about the finite-sample performance of estimation and model selection. We also contribute to the theory of clustering with an over-specified number of clusters and derive new convergence rates for this setting. Our results suggest that over-fitting can be severe in k-means style estimators when the number of clusters is over-specified.

econ.EM