SearcharxivSearch

arXiv subjects

Ben Deaner

Publications and source records attributed to Ben Deaner.

11 recordsLinked to original sources

Linear Estimation of Structural and Causal Effects for Nonseparable Panel Data

This paper develops linear estimators for structural and causal parameters of nonseparable models using panel data. These models incorporate unobserved, time-varying, individual heterogeneity, which may be correlated with the regressors. Estimation is based on an approximation of a conditional average potential outcome by a linear sieve specification with individual-specific parameters. Effects of interest are estimated by a bias corrected average of individual ridge regressions. We demonstrate how this approach can be applied to estimate causal effects, counterfactual consumer welfare, and averages of individual taxable income elasticities. We show that the proposed estimator has an empirical Bayes interpretation and possesses a number of other useful properties. We formulate Large-$T$ asymptotics that can accommodate discrete regressors and which bypass partial identification in this case. We employ the methods to estimate average equivalent variation and deadweight loss for potential price increases using data on grocery purchases.

econ.EM

When unlearning is free: leveraging low influence points to reduce computational costs

As concerns around data privacy in machine learning grow, the ability to unlearn, or remove, specific data points from trained models becomes increasingly important. While state of the art unlearning methods have emerged in response, they typically treat all points in the forget set equally. In this work, we challenge this approach by asking whether points that have a negligible impact on the model's learning need to be removed. Through a comparative analysis of influence functions across language and vision tasks, we identify subsets of training data with negligible impact on model outputs. Leveraging this insight, we propose an efficient unlearning framework that reduces the size of datasets before unlearning leading to significant computational savings (up to approximately 50 percent) on real world empirical examples.

cs.LG

Extrapolation in Regression Discontinuity Design Using Comonotonicity

We present a novel approach for extrapolating causal effects away from the margin between treatment and non-treatment in sharp regression discontinuity designs with multiple covariates. Our methods apply both to settings in which treatment is a function of multiple observables and settings in which treatment is determined based on a single running variable. Our key identifying assumption is that conditional average treated and untreated potential outcomes are comonotonic: covariate values associated with higher average untreated potential outcomes are also associated with higher average treated potential outcomes. We provide an estimation method based on local linear regression. Our estimands are weighted average causal effects, even if comonotonicity fails. We apply our methods to evaluate counterfactual mandatory summer school policies.

econ.EM

Inferring Treatment Effects in Large Panels by Uncovering Latent Similarities

The presence of unobserved confounders is one of the main challenges in identifying treatment effects. In this paper, we propose a new approach to causal inference using panel data with large large $N$ and $T$. Our approach imputes the untreated potential outcomes for treated units using the outcomes for untreated individuals with similar values of the latent confounders. In order to find units with similar latent characteristics, we utilize long pre-treatment histories of the outcomes. Our analysis is based on a nonparametric, nonlinear, and nonseparable factor model for untreated potential outcomes and treatments. The model satisfies minimal smoothness requirements. We impute both missing counterfactual outcomes and propensity scores using kernel smoothing based on the constructed measure of latent similarity between units, and demonstrate that our estimates can achieve the optimal nonparametric rate of convergence up to log terms. Using these estimates, we construct a doubly robust estimator of the period-specifc average treatment effect on the treated (ATT), and provide conditions, under which this estimator is $\sqrt{N}$-consistent, and asymptotically normal and unbiased. Our simulation study demonstrates that our method provides accurate inference for a wide range of data generating processes.

econ.EM

Density Ratio-based Proxy Causal Learning Without Density Ratios

We address the setting of Proxy Causal Learning (PCL), which has the goal of estimating causal effects from observed data in the presence of hidden confounding. Proxy methods accomplish this task using two proxy variables related to the latent confounder: a treatment proxy (related to the treatment) and an outcome proxy (related to the outcome). Two approaches have been proposed to perform causal effect estimation given proxy variables; however only one of these has found mainstream acceptance, since the other was understood to require density ratio estimation - a challenging task in high dimensions. In the present work, we propose a practical and effective implementation of the second approach, which bypasses explicit density ratio estimation and is suitable for continuous and high-dimensional treatments. We employ kernel ridge regression to derive estimators, resulting in simple closed-form solutions for dose-response and conditional dose-response curves, along with consistency guarantees. Our methods empirically demonstrate superior or comparable performance to existing frameworks on synthetic and real-world datasets.

cs.LG

Causal Duration Analysis with Diff-in-Diff

In economic program evaluation, it is common to obtain panel data in which outcomes are indicators that an individual has reached an absorbing state. For example, they may indicate whether an individual has exited a period of unemployment, passed an exam, left a marriage, or had their parole revoked. The parallel trends assumption that underpins difference-in-differences generally fails in such settings. We suggest identifying conditions similar to those of difference-in-differences, but which apply to hazard rates rather than mean outcomes. These alternative assumptions motivate estimators that retain the simplicity and transparency of standard diff-in-diff, and corresponding specification tests. Our approach can be adapted to include general linear restrictions between the hazard rates of different groups, motivating duration analogues of the triple differences and synthetic control methods. We apply our procedures to examine the impact of a policy that increased the generosity of unemployment benefits, using a cross-cohort comparison.

econ.EM

Controlling for Latent Confounding with Triple Proxies

We present new results for nonparametric identification of causal effects using noisy proxies for unobserved confounders. Our approach builds on the results of \citet{Hu2008} who tackle the problem of general measurement error. We call this the `triple proxy' approach because it requires three proxies that are jointly independent conditional on unobservables. We consider three different choices for the third proxy: it may be an outcome, a vector of treatments, or a collection of auxiliary variables. We compare to an alternative identification strategy introduced by \citet{Miao2018a} in which causal effects are identified using two conditionally independent proxies. We refer to this as the `double proxy' approach. The triple proxy approach identifies objects that are not identified by the double proxy approach, including some that capture the variation in average treatment effects between strata of the unobservables. Moreover, the conditional independence assumptions in the double and triple proxy approaches are non-nested.

econ.EM

Many Proxy Controls

A recent literature considers causal inference using noisy proxies for unobserved confounding factors. The proxies are divided into two sets that are independent conditional on the confounders. One set of proxies are `negative control treatments' and the other are `negative control outcomes'. Existing work applies to low-dimensional settings with a fixed number of proxies and confounders. In this work we consider linear models with many proxy controls and possibly many confounders. A key insight is that if each group of proxies is strictly larger than the number of confounding factors, then a matrix of nuisance parameters has a low-rank structure and a vector of nuisance parameters has a sparse structure. We can exploit the rank-restriction and sparsity to reduce the number of free parameters to be estimated. The number of unobserved confounders is not known a priori but we show that it is identified, and we apply penalization methods to adapt to this quantity. We provide an estimator with a closed-form as well as a doubly-robust estimator that must be evaluated using numerical methods. We provide conditions under which our doubly-robust estimator is uniformly root-$n$ consistent, asymptotically centered normal, and our suggested confidence intervals have asymptotically correct coverage. We provide simulation evidence that our methods achieve better performance than existing approaches in high dimensions, particularly when the number of proxies is substantially larger than the number of confounders.

econ.EM

Approximation-Robust Inference in Dynamic Discrete Choice

Estimation and inference in dynamic discrete choice models often relies on approximation to lower the computational burden of dynamic programming. Unfortunately, the use of approximation can impart substantial bias in estimation and results in invalid confidence sets. We present a method for set estimation and inference that explicitly accounts for the use of approximation and is thus valid regardless of the approximation error. We show how one can account for the error from approximation at low computational cost. Our methodology allows researchers to assess the estimation error due to the use of approximation and thus more effectively manage the trade-off between bias and computational expedience. We provide simulation evidence to demonstrate the practicality of our approach.

econ.EM

Nonparametric Instrumental Variables Estimation Under Misspecification

Nonparametric Instrumental Variables (NPIV) analysis is based on a conditional moment restriction. We show that if this moment condition is even slightly misspecified, say because instruments are not quite valid, then NPIV estimates can be subject to substantial asymptotic error and the identified set under a relaxed moment condition may be large. Imposing strong a priori smoothness restrictions mitigates the problem but induces bias if the restrictions are too strong. In order to manage this trade-off we develop a methods for empirical sensitivity analysis and apply them to the consumer demand data previously analyzed in Blundell (2007) and Horowitz (2011).

econ.EM

Proxy Controls and Panel Data

We provide new results for nonparametric identification, estimation, and inference of causal effects using `proxy controls': observables that are noisy but informative proxies for unobserved confounding factors. Our analysis applies to cross-sectional settings but is particularly well-suited to panel models. Our identification results motivate a simple and `well-posed' nonparametric estimator. We derive convergence rates for the estimator and construct uniform confidence bands with asymptotically correct size. In panel settings, our methods provide a novel approach to the difficult problem of identification with non-separable, general heterogeneity and fixed $T$. In panels, observations from different periods serve as proxies for unobserved heterogeneity and our key identifying assumptions follow from restrictions on the serial dependence structure. We apply our methods to two empirical settings. We estimate consumer demand counterfactuals using panel data and we estimate causal effects of grade retention on cognitive performance.

econ.EM