SearcharxivSearch

arXiv subjects

Christoph Rothe

Publications and source records attributed to Christoph Rothe.

8 recordsLinked to original sources

Inference in Regression Discontinuity Designs with Clustered Data

Clustered sampling is prevalent in empirical regression discontinuity (RD) designs, but it has not received much attention in the theoretical literature. In this paper, we introduce a general model-based framework for such settings and derive high-level conditions under which the standard local linear RD estimator is asymptotically normal. We verify that our high-level assumptions hold across a wide range of empirical designs, including settings of growing cluster sizes. We further show that clustered standard errors that are currently used in practice can be either inconsistent or overly conservative in finite samples. To address these issues, we propose a novel nearest-neighbor-type variance estimator and illustrate its properties in a diverse set of empirical applications.

econ.EM

Bias-Aware Inference in Fuzzy Regression Discontinuity Designs

We propose new confidence sets (CSs) for the regression discontinuity parameter in fuzzy designs. Our CSs are based on local linear regression, and are bias-aware, in the sense that they take possible bias explicitly into account. Their construction shares similarities with that of Anderson-Rubin CSs in exactly identified instrumental variable models, and thereby avoids issues with "delta method" approximations that underlie most commonly used existing inference methods for fuzzy regression discontinuity analysis. Our CSs are asymptotically equivalent to existing procedures in canonical settings with strong identification and a continuous running variable. However, due to their particular construction they are also valid under a wide range of empirically relevant conditions in which existing methods can fail, such as setups with discrete running variables, donut designs, and weak identification.

econ.EM

Inference in Regression Discontinuity Designs with High-Dimensional Covariates

We study regression discontinuity designs in which many predetermined covariates, possibly much more than the number of observations, can be used to increase the precision of treatment effect estimates. We consider a two-step estimator which first selects a small number of "important" covariates through a localized Lasso-type procedure, and then, in a second step, estimates the treatment effect by including the selected covariates linearly into the usual local linear estimator. We provide an in-depth analysis of the algorithm's theoretical properties, showing that, under an approximate sparsity condition, the resulting estimator is asymptotically normal, with asymptotic bias and variance that are conceptually similar to those obtained in low-dimensional settings. Bandwidth selection and inference can be carried out using standard methods. We also provide simulations and an empirical application.

econ.EM

Flexible Covariate Adjustments in Regression Discontinuity Designs

Empirical regression discontinuity (RD) studies often include covariates in their specifications to increase the precision of their estimates. In this paper, we propose a novel class of estimators that use such covariate information more efficiently than existing methods and can accommodate many covariates. Our estimators are simple to implement and involve running a standard RD analysis after subtracting a function of the covariates from the original outcome variable. We characterize the function of the covariates that minimizes the asymptotic variance of these estimators. We also show that the conventional RD framework gives rise to a special robustness property which implies that the optimal adjustment function can be estimated flexibly via modern machine learning techniques without affecting the first-order properties of the final RD estimator. We demonstrate our methods' scope for efficiency improvements by reanalyzing data from a large number of recently published empirical studies.

econ.EM

Combining Population and Study Data for Inference on Event Rates

This note considers the problem of conducting statistical inference on the share of individuals in some subgroup of a population that experience some event. The specific complication is that the size of the subgroup needs to be estimated, whereas the number of individuals that experience the event is known. The problem is motivated by the recent study of Streeck et al. (2020), who estimate the infection fatality rate (IFR) of SARS-CoV-2 infection in a German town that experienced a super-spreading event in mid-February 2020. In their case the subgroup of interest is comprised of all infected individuals, and the event is death caused by the infection. We clarify issues with the precise definition of the target parameter in this context, and propose confidence intervals (CIs) based on classical statistical principles that result in good coverage properties.

stat.AP

Inference in Regression Discontinuity Designs with a Discrete Running Variable

We consider inference in regression discontinuity designs when the running variable only takes a moderate number of distinct values. In particular, we study the common practice of using confidence intervals (CIs) based on standard errors that are clustered by the running variable as a means to make inference robust to model misspecification (Lee and Card, 2008). We derive theoretical results and present simulation and empirical evidence showing that these CIs do not guard against model misspecification, and that they have poor coverage properties. We therefore recommend against using these CIs in practice. We instead propose two alternative CIs with guaranteed coverage properties under easily interpretable restrictions on the conditional expectation function.

stat.AP

Estimating Derivatives of Function-Valued Parameters in a Class of Moment Condition Models

We develop a general approach to estimating the derivative of a function-valued parameter $θ_o(u)$ that is identified for every value of $u$ as the solution to a moment condition. This setup in particular covers many interesting models for conditional distributions, such as quantile regression or distribution regression. Exploiting that $θ_o(u)$ solves a moment condition, we obtain an explicit expression for its derivative from the Implicit Function Theorem, and estimate the components of this expression by suitable sample analogues, which requires the use of (local linear) smoothing. Our estimator can then be used for a variety of purposes, including the estimation of conditional density functions, quantile partial effects, and structural auction models in economics.

stat.ME

Nonparametric regression with nonparametrically generated covariates

We analyze the statistical properties of nonparametric regression estimators using covariates which are not directly observable, but have be estimated from data in a preliminary step. These so-called generated covariates appear in numerous applications, including two-stage nonparametric regression, estimation of simultaneous equation models or censored regression models. Yet so far there seems to be no general theory for their impact on the final estimator's statistical properties. Our paper provides such results. We derive a stochastic expansion that characterizes the influence of the generation step on the final estimator, and use it to derive rates of consistency and asymptotic distributions accounting for the presence of generated covariates.

math.ST