SearcharxivSearch

arXiv subjects

Ben B. Hansen

Publications and source records attributed to Ben B. Hansen.

8 recordsLinked to original sources

Propensity score adjustment when errors in achievement measures inform treatment assignment

U.S. state education agencies mark schools displaying achievement gaps between demographic subgroups as needing improvement. Some schools may have few students in these subgroups, such that average end-of-year test scores only noisily measure the average "true" score-the score one would expect if students took the test many times. This, in addition to the masking of small subgroup averages in publicly available assessment data, poses challenges for evaluating interventions aimed at closing achievement gaps. We introduce propensity score estimates designed to achieve balance on subgroup average true scores. These estimates are available even when noisy measurements are not and improve overlap compared to those that ignore measurement error, leading to greater bias reduction of matching estimators. We demonstrate our methods through simulation and an application to a statewide initiative in Texas for curbing summer learning loss.

stat.ME

Undersmoothed LASSO Models for Propensity Score Weighting and Synthetic Negative Control Exposures for Bias Detection

The propensity score (PS) is often used to control for large numbers of covariates in high-dimensional healthcare database studies. The least absolute shrinkage and selection operator (LASSO) has become the most widely used tool for fitting large-scale PS models in these settings. LASSO uses L1 regularized regression to prevent overfitting by shrinking coefficients toward zero (setting some exactly to zero). The degree of regularization is typically selected using cross-validation to minimize out-of-sample prediction error. Both theory and simulations have shown, however, that when using LASSO models for PS weighting, less regularization is needed to minimize bias in PS weighted estimators. This is referred to as undersmoothing the LASSO model, where the optimal degree of undersmoothing can be derived from the target causal parameter's efficient influence function. In many settings, however, the efficient influence function is unknown or difficult to derive. Here, we consider the use of balance metrics as a simple and generally applicable approach to select the degree of undersmoothing when the efficient influence function is unknown. Because LASSO models that are tuned using balance metrics alone are not assured to minimize bias in PS weighted estimators -- as such metrics are blind to the efficient influence function -- we propose a framework to generate synthetic negative control exposures for bias detection. We show that synthetic negative control exposures can identify analyses that likely violate partial exchangeability due to lack of control for measured confounding. Finally, we use a series of numerical studies to investigate the finite sample performance of using balance criteria to undersmooth LASSO PS-weighted estimators, and the use of synthetic negative control exposures to detect biased analyses.

stat.ME

Robust Design-Based Estimation and Inference for Stratified Randomized Trials with Varying Cluster Sizes

Clustered randomized controlled trials are often stratified or pair-matched to improve covariate balance and efficiency. Sample average treatment effects (SATEs) are commonly estimated by averaging stratum-level treatment-control mean contrasts -- an approach that is natural and widely used. We show that, in stratified clustered trials with heterogeneous cluster sizes, such estimators need not be consistent for the SATE. They can converge to the wrong limit even under correct randomization and without model misspecification. The source is a covariance between cluster sizes and treatment effects: stratumwise averaging mis-weights clusters in a way that produces bias of constant order, regardless of sample size. We study the H\'ajek (ratio) estimator as a robust alternative. By aggregating outcomes within treatment groups before taking their difference, it remains consistent in clustered trials that grow by increasing strata sizes or the number of strata. Despite that, its use in design-based analyses of clustered trials has been limited by the lack of variance estimators. We develop a design-based variance estimator that applies to any number of strata of any size, and show that it is asymptotically conservative, a property that holds even when some strata contain only a single treated or control unit. We also present tests improving the coverage of Wald tests when the number of clusters is moderate. The framework extends naturally to covariate-adjusted estimators via a variance orthogonality property.

stat.ME

Matching calipers and the precision of index estimation

This paper characterizes the precision of index estimation as it carries over into precision of matching. In a model assuming Gaussian covariates and making best-case assumptions about matching quality, it sharply characterizes average and worst-case discrepancies between paired differences of true versus estimated index values. In this optimistic setting, worst-case true and estimated index differences decline to zero if $p=o[n/(\log n)]$, the same restriction on model size that is needed for consistency of common index models. This remains so as the Gaussian assumption is relaxed to sub-gaussian, if in that case the characterization of paired index errors is less sharp. The formula derived under Gaussian assumptions is used as the basis for a matching caliper. Matching such that paired differences on the estimated index fall below this caliper brings the benefit that after matching, worst-case differences onan underlying index tend to 0 if $p = o\{[n/(\log n)]^{2/3}\}$. (With a linear index model, $p=o[n/(\log n)]$ suffices.) A proposed refinement of the caliper condition brings the same benefits without the sub-gaussian condition on covariates. When strong ignorability holds and the index is a well-specified propensity or prognostic score, ensuring in this way that worst-case matched discrepancies on it tend to 0 with increasing $n$ also ensures the consistency of matched estimators of the treatment effect.

stat.ME

An Aggregation Scheme for Increased Power

We present an aggregation scheme that increases power in randomized controlled trials and quasi-experiments when the intervention possesses a robust and well-articulated theory of change. Longitudinal data analyzing interventions often include multiple observations on individuals, some of which may be more likely to manifest a treatment effect than others. An intervention's theory of change provides guidance as to which of those observations are best situated to exhibit that treatment effect. Our power-maximizing weighting for repeated-measurements with delayed-effects scheme, PWRD aggregation, converts the theory of change into a test statistic with improved asymptotic relative efficiency, delivering tests with greater statistical power. We illustrate this method on an IES-funded cluster randomized trial testing the efficacy of a reading intervention designed to assist early elementary students at risk of falling behind their peers. The salient theory of change holds program benefits to be delayed and non-uniform, experienced after a student's performance stalls. In this instance, the PWRD technique's effect on power is found to be comparable to that of doubling the number of clusters in the experiment.

stat.ME

Limitless Regression Discontinuity

Conventionally, regression discontinuity analysis contrasts a univariate regression's limits as its independent variable, $R$, approaches a cut-point, $c$, from either side. Alternative methods target the average treatment effect in a small region around $c$, at the cost of an assumption that treatment assignment, $\mathcal{I}\left[R<c\right]$, is ignorable vis a vis potential outcomes. Instead, the method presented in this paper assumes Residual Ignorability, ignorability of treatment assignment vis a vis detrended potential outcomes. Detrending is effected not with ordinary least squares but with MM-estimation, following a distinct phase of sample decontamination. The method's inferences acknowledge uncertainty in both of these adjustments, despite its applicability whether $R$ is discrete or continuous; it is uniquely robust to leading validity threats facing regression discontinuity designs.

stat.AP

The sensitivity of linear regression coefficients' confidence limits to the omission of a confounder

Omitted variable bias can affect treatment effect estimates obtained from observational data due to the lack of random assignment to treatment groups. Sensitivity analyses adjust these estimates to quantify the impact of potential omitted variables. This paper presents methods of sensitivity analysis to adjust interval estimates of treatment effect---both the point estimate and standard error---obtained using multiple linear regression. Central to our approach is what we term benchmarking, the use of data to establish reference points for speculation about omitted confounders. The method adapts to treatment effects that may differ by subgroup, to scenarios involving omission of multiple variables, and to combinations of covariance adjustment with propensity score stratification. We illustrate it using data from an influential study of health outcomes of patients admitted to critical care.

stat.ME

Covariate Balance in Simple, Stratified and Clustered Comparative Studies

In randomized experiments, treatment and control groups should be roughly the same--balanced--in their distributions of pretreatment variables. But how nearly so? Can descriptive comparisons meaningfully be paired with significance tests? If so, should there be several such tests, one for each pretreatment variable, or should there be a single, omnibus test? Could such a test be engineered to give easily computed $p$-values that are reliable in samples of moderate size, or would simulation be needed for reliable calibration? What new concerns are introduced by random assignment of clusters? Which tests of balance would be optimal? To address these questions, Fisher's randomization inference is applied to the question of balance. Its application suggests the reversal of published conclusions about two studies, one clinical and the other a field experiment in political participation.

stat.ME