SearcharxivSearch

arXiv subjects

Edward H. Kennedy

Publications and source records attributed to Edward H. Kennedy.

At least 19 recordsLinked to original sources

On regression with estimated covariates and conditional effects given the propensity score

Motivated by the study of heterogeneous returns to education in Brand & Xie 2010, which considers how the effect of completing college on earnings varies with the (unknown) probability of completing college, we analyze the problem of estimating a nonparametric regression function when certain covariates are estimated in a first step. Plug-in estimators that treat the estimated covariates as known generally suffer from first-stage estimation error. To mitigate this issue, we analyze two debiasing approaches within a framework that is agnostic to the choice of the first-stage estimation method and relies on either local-smoothing or sieve-based methods for the second-stage regression. In particular, we consider: (i) influence function-based estimators of pathwise differentiable parameters that approximate the target estimand, and (ii) a variant of plug-in estimators that directly aims to correct their bias. For each method, we upper bound the estimation error and characterize conditions under which oracle rates can be approached, highlighting the possible gains in terms of convergence rates relative to the plug-ins. Simulation studies illustrate the finite-sample behavior of the methods. We apply our methodology to data from the National Longitudinal Survey of Youth 1997 and find evidence that completing college yields the largest reductions in unemployment for individuals least likely to do so, consistent with earlier findings in the literature (Brand & Xie 2010; Brand 2023).

stat.ME

Causal Inference with High-Dimensional Treatments

In this work, we consider causal inference in various high-dimensional treatment settings, including for single multi-valued treatments and vector treatments with binary or continuous components, when the number of treatments can be comparable to or even larger than the number of observations. These settings bring unique challenges: first, the treatment effects of interest are represented by a high-dimensional vector rather than a scalar; second, positivity violations are often unavoidable; and third, estimation can be based on a smaller effective sample size. We first discuss fundamental limits of estimating effects here, showing that consistent estimation is impossible without further assumptions. We go on to propose novel doubly robust estimators for mean potential outcomes of a high-dimensional single multi-valued treatment. We analyze the proposed estimators under sparsity assumptions, giving finite-sample risk bounds and showing that consistent estimation is possible under these conditions. Moreover, we derive minimax lower bounds in a sparse and structure-agnostic model to characterize optimal rates of convergence and show our risk bounds are unimprovable. We then generalize our proposed estimators as a sparse pseudo-outcome regression framework with constrained regression estimators and error guarantees under sparsity, allowing estimation of generic functionals and different types of high-dimensional treatments. We apply the framework to derive estimators of the mean potential outcomes for high-dimensional vector treatments. Finally, we illustrate the proposed methods through a simulation and an empirical application.

math.ST

On the Equivalence between Neyman Orthogonality and Pathwise Differentiability

It has been frequently observed that Neyman orthogonality, the central device underlying double/debiased machine learning (Chernozhukov et al., 2018), and pathwise differentiability, a cornerstone concept from semiparametric theory, often lead to the same debiased estimators in practice. Despite the widespread adoption of both ideas, the precise nature of this equivalence has remained elusive, with the two concepts having been developed in largely separate traditions. In this work, we revisit the semiparametric framework of van der Laan and Robins (2003) and identify an implicit regularity assumption on the relationship between target and nuisance parameters -- a local product structure -- that allows us to establish a formal equivalence between Neyman orthogonality and pathwise differentiability. We also show that the two directions of this equivalence impose fundamentally different structural requirements. Finally, we illustrate the theory through three detailed examples of estimating the average treatment effect and expected density in a nonparametric model, as well as the slope in a partially linear model. This helps clarify the relationship between these two foundational frameworks and provides a useful reference for practitioners working at their intersection.

stat.ME

Causal Inference with High-dimensional Discrete Covariates

When estimating causal effects from observational studies, researchers often need to adjust for many covariates to deconfound the non-causal relationship between exposure and outcome, among which many covariates are discrete. The behavior of commonly used estimators in the presence of many discrete covariates is not well understood since their properties are often analyzed under structural assumptions including sparsity and smoothness, which do not apply in discrete settings. In this work, we study the estimation of causal effects in a model where the covariates required for confounding adjustment are discrete but high-dimensional, meaning the number of categories $d$ is comparable with or even larger than sample size $n$. Specifically, we show the mean squared error of commonly used regression, weighting and doubly robust estimators is bounded by $\frac{d^2}{n^2}+\frac{1}{n}$. We then prove the minimax lower bound for the average treatment effect is of order $\frac{d^2}{n^2 \log^2 n}+\frac{1}{n}$, which characterizes the fundamental difficulty of causal effect estimation in the high-dimensional discrete setting, and shows the estimators mentioned above are rate-optimal up to log-factors. We further consider additional structures that can be exploited, namely effect homogeneity and prior knowledge of the covariate distribution, and propose new estimators that enjoy faster convergence rates of order $\frac{d}{n^2} + \frac{1}{n}$, which achieve consistency in a broader regime. The results are illustrated empirically via simulation studies.

math.ST

Fast convergence rates for dose-response estimation

We consider the problem of estimating a dose-response curve. Continuous treatments arise often in practice, e.g. in the form of time spent on an operation, distance traveled to a location or dosage of a drug. Letting $A$ denote a continuous treatment variable, the target of inference is the expected outcome if everyone in the population takes treatment level $A=t$. Under standard assumptions, the dose-response function takes the form of a partial mean. Building upon the recent literature on nonparametric regression with estimated outcomes, our first contribution is to study global and local estimators of the dose-response based on empirical risk minimization. Our second and main contribution is to construct a $m^{\text{th}}$-order estimator based on the theory of higher-order influence functions. Under certain conditions, this higher order estimator achieves the fastest rate of convergence that we are aware of for this problem. However, the other two approaches are easier to implement using off-the-shelf software, since they are formulated as two-stage regression tasks. For each estimator, we provide an upper bound on the mean-square error and investigate its finite-sample performance through simulations and an empirical application. Finally, the supplementary material introduces a flexible, nonparametric approach for sensitivity analysis to violations of the no-unmeasured-confounding assumption with continuous treatments.

stat.ME

Causal K-Means Clustering

Causal effects are often characterized with population summaries. These might provide an incomplete picture when there are heterogeneous treatment effects across subgroups. Since the subgroup structure is typically unknown, it is more challenging to identify and evaluate subgroup effects than population effects. We propose a new solution to this problem: \emph{Causal k-Means Clustering}, which leverages the k-means clustering algorithm to uncover the unknown subgroup structure. Our problem differs significantly from the conventional clustering setup since the variables to be clustered are unknown counterfactual functions. We present a plug-in estimator which is simple and readily implementable using off-the-shelf algorithms, and study its rate of convergence. We also develop a new bias-corrected estimator based on nonparametric efficiency theory and double machine learning, and show that this estimator achieves fast root-n rates and asymptotic normality in large nonparametric models. Our proposed methods are especially useful for modern outcome-wide studies with multiple treatment levels. Further, our framework is extensible to clustering with generic pseudo-outcomes, such as partially observed outcomes or otherwise unknown functions. Finally, we explore finite sample properties via simulation, and illustrate the proposed methods using a study of mobile-supported self-management for chronic low back pain.

stat.ME

Doubly-Robust Functional Average Treatment Effect Estimation

Understanding causal relationships in the presence of complex, structured data remains a central challenge in modern statistics and science in general. While traditional causal inference methods are well-suited for scalar outcomes, many scientific applications demand tools capable of handling functional data -- outcomes observed as functions over continuous domains such as time or space. Motivated by this need, we propose DR-FoS, a novel method for estimating the Functional Average Treatment Effect (FATE) in observational studies with functional outcomes. DR-FoS exhibits double robustness properties, ensuring consistent estimation of FATE even if either the outcome or the treatment assignment model is misspecified. By leveraging recent advances in functional data analysis and causal inference, we establish the asymptotic properties of the estimator, proving its convergence to a Gaussian process. This guarantees valid inference with simultaneous confidence bands across the entire functional domain. Through extensive simulations, we show that DR-FoS achieves robust performance under a wide range of model specifications. Finally, we illustrate the utility of DR-FoS in a real-world application, analyzing functional outcomes to uncover meaningful causal insights in the SHARE ({\em Survey of Health, Aging and Retirement in Europe}) dataset.

stat.ME

Doubly Robust Machine Learning for Population Size Estimation with Missing Covariates: Application to Gaza Conflict Mortality

Population size estimation from capture-recapture data is central for studying hard-to-reach populations, incorporating auxiliary covariates to account for heterogeneous capture probabilities and recapture dependencies. However, missing attributes pose a critical methodological challenge due to reluctance to share sensitive information, data collection limitations, and imperfect record linkage. Existing approaches either ignore missingness or rely on a priori imputation, potentially introducing substantial bias. In this work, we develop a novel nonparametric estimation framework using a Missing at Random assumption to identify capture probabilities under missing covariates. Using semiparametric efficiency theory, we construct one-step estimators that combine efficiency, robustness, and finite-sample validity: they approximately achieve the nonparametric efficiency bound, accommodate flexible machine learning methods through a doubly robust structure, and provide approximately valid inference for any sample size. Simulations demonstrate substantial improvements over naive imputation approaches, with our doubly robust ML estimators maintaining valid inference even at high missingness rates where competing methods fail. We apply our methodology to re-estimate mortality in the Gaza Strip from October 7, 2023, to June 30, 2024, using three-list capture-recapture data with missing demographic information. Our approach yields more conservative yet precise estimates compared to previous methods, indicating the true death toll exceeds official statistics by approximately 26%. Our framework provides practitioners with principled tools for handling incomplete data in conflict settings and other applications with hard-to-reach populations.

stat.ME

Incremental effects for continuous exposures

Causal inference problems often involve continuous treatments, such as dose, duration, or frequency. However, identifying and estimating standard dose-response estimands requires that everyone has some chance of receiving any level of the exposure (i.e., positivity). To avoid this assumption, we consider stochastic interventions based on exponentially tilting the treatment distribution by some parameter $δ$ (an incremental effect); this increases or decreases the likelihood a unit receives a given treatment level. We derive the efficient influence function and semiparametric efficiency bound for these incremental effects under continuous exposures. We then show estimation depends on the size of the tilt, as measured by $δ$. In particular, we derive new minimax lower bounds illustrating how the best possible root mean squared error scales with an effective sample size of $n / δ$, instead of $n$. Further, we establish new convergence rates and bounds on the bias of double machine learning-style estimators. Our novel analysis gives a better dependence on $δ$ compared to standard analyses by using mixed supremum and $L_2$ norms. Finally, we define a "reflected" exponential tilt around any interior point and show that taking $δ\to \infty$ yields a new estimator of the dose-response curve across the treatment support.

stat.ME

Sensitivity analysis for incremental effects, with application to a study of victimization & offending

Sensitivity analysis for unmeasured confounding under incremental propensity score interventions remains relatively underdeveloped. Incremental interventions define stochastic treatment regimes by multiplying the odds of treatment, offering a flexible framework for causal effect estimation. To study incremental effects when there are unobserved confounders, we adopt Rosenbaum's sensitivity model in single time point settings, and propose a doubly robust estimator for the resulting effect bounds. The bound estimators are asymptotically normal under mild conditions on nuisance function estimation. We show that incremental effect bounds can be narrower or wider than those for mean potential outcomes, and that the bounds must lie between the expected minimum and maximum of the conditional bounds on E(Y^0|X) and E(Y^1|X). For time-varying treatments, we consider the marginal sensitivity model. Although sharp bounds for incremental effects are identifiable from longitudinal data under this model, practical estimators have not yet been established; we discuss this challenge and provide partial results toward implementation. Finally, we apply our methods to study the effect of victimization on subsequent offending using data from the National Longitudinal Study of Adolescent to Adult Health (Add Health), illustrating the robustness of our findings in an empirical setting.

stat.ME

Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random

The Difference-in-Differences (DiD) method is a fundamental tool for causal inference, yet its application is often complicated by missing data. Although recent work has developed robust DiD estimators for complex settings like staggered treatment adoption, these methods typically assume complete data and fail to address the critical challenge of outcomes that are missing at random (MAR) -- a common problem that invalidates standard estimators. We develop a rigorous framework, rooted in semiparametric theory, for identifying and efficiently estimating the Average Treatment Effect on the Treated (ATT) when either pre- or post-treatment (or both) outcomes are missing at random. We first establish nonparametric identification of the ATT under two minimal sets of sufficient conditions. For each, we derive the semiparametric efficiency bound, which provides a formal benchmark for asymptotic optimality. We then propose novel estimators that are asymptotically efficient, achieving this theoretical bound. A key feature of our estimators is their multiple robustness, which ensures consistency even if some nuisance function models are misspecified. We validate the properties of our estimators and showcase their broad applicability through an extensive simulation study.

stat.ME

Distribution-uniform anytime-valid sequential inference and the Robbins-Siegmund distributions

This paper develops a theory of distribution- and time-uniform asymptotics, culminating in the first large-sample anytime-valid inference procedures that are shown to be uniformly valid in a rich class of distributions. Historically, anytime-valid methods -- including confidence sequences, anytime $p$-values, and sequential hypothesis tests -- have been justified nonasymptotically. By contrast, large-sample inference procedures such as those based on the central limit theorem occupy an important part of statistical toolbox due to their simplicity, universality, and the weak assumptions they make. While recent work has derived asymptotic analogues of anytime-valid methods, they were not distribution-uniform (also called \emph{honest}), meaning that their type-I errors may not be uniformly upper-bounded by the desired level in the limit. The theory and methods we outline resolve this tension, and they do so without imposing assumptions that are any stronger than the distribution-uniform fixed-$n$ (non-anytime-valid) counterparts or distribution-pointwise anytime-valid special cases. It is shown that certain ``Robbins-Siegmund'' probability distributions play roles in anytime-valid asymptotics analogous to those played by Gaussian distributions in standard asymptotics. As an application, we derive the first anytime-valid test of conditional independence without the Model-X assumption.

math.ST

Handling Missing Responses under Cluster Dependence with Applications to Language Model Evaluation

Human annotations play a crucial role in evaluating the performance of GenAI models. Two common challenges in practice, however, are missing annotations (the response variable of interest) and cluster dependence among human-AI interactions (e.g., questions asked by the same user may be highly correlated). Reliable inference must address both these issues to achieve unbiased estimation and appropriately quantify uncertainty when estimating average scores from human annotations. In this paper, we analyze the doubly robust estimator, a widely used method in missing data analysis and causal inference, applied to this setting and establish novel theoretical properties under cluster dependence. We further illustrate our findings through simulations and a real-world conversation quality dataset. Our theoretical and empirical results underscore the importance of incorporating cluster dependence in missing response problems to perform valid statistical inference.

stat.ME

Semiparametric sensitivity analysis: unmeasured confounding in observational studies

Establishing cause-effect relationships from observational data often relies on untestable assumptions. It is crucial to know whether, and to what extent, the conclusions drawn from non-experimental studies are robust to potential unmeasured confounding. In this paper, we focus on the average causal effect (ACE) as our target of inference. We generalize the sensitivity analysis approach developed by Robins et al. (2000), Franks et al. (2020), and Zhou and Yao (2023). We use semiparametric theory to derive the non-parametric efficient influence function of the ACE, for fixed sensitivity parameters. We use this influence function to construct a one-step, split sample, truncated estimator of the ACE. Our estimator depends on semiparametric models for the distribution of the observed data; importantly, these models do not impose any restrictions on the values of sensitivity analysis parameters. We establish sufficient conditions ensuring that our estimator has root-n asymptotics. We use our methodology to evaluate the causal effect of smoking during pregnancy on birth weight. We also evaluate the performance of estimation procedure in a simulation study.

stat.ME

Calibrated sensitivity models

In causal inference, sensitivity models assess how unmeasured confounders could alter causal analyses, but the sensitivity parameter -- which quantifies the degree of unmeasured confounding -- is often difficult to interpret. For this reason, researchers sometimes compare the sensitivity parameter to an estimate of measured confounding. This is known as calibration, or benchmarking. However, calibrated estimates are not always interpreted correctly, and uncertainty in the estimate of measured confounding is rarely accounted for. To address these limitations, we propose calibrated sensitivity models, which directly bound the degree of unmeasured confounding by a multiple of measured confounding. We develop a clear framework for interpreting calibrated sensitivity models and derive statistical methods for accounting for uncertainty due to estimating measured confounding. Incorporating this uncertainty shows causal analyses may be either less or more robust to unmeasured confounding than suggested by standard approaches. We develop efficient estimators and inferential methods for bounds on the average treatment effect with three calibrated sensitivity models, establishing parametric efficiency and asymptotic normality under doubly robust style nonparametric conditions. We illustrate our methods with an analysis of the effect of mothers' smoking on infant birthweight.

stat.ME

Nonparametric Estimation of Local Treatment Effects with Continuous Instruments

Instrumental variable methods are widely used to address unmeasured confounding, yet much of the existing literature has focused on the binary instrument setting. Extensions to continuous instruments often impose strong parametric assumptions for identification and estimation, which can be difficult to justify and may limit their applicability in complex real-world settings. In this work, we develop theory and methods for nonparametric estimation of treatment effects with a continuous instrumental variable. We introduce an estimand that, under a monotonicity assumption, quantifies the treatment effect among the maximal complier class, generalizing the local average treatment effect framework to continuous instruments. Considering this estimand and the local instrumental variable curve, we draw connections to the dose-response function and its derivative, and propose doubly robust estimation methods. We establish convergence rates and conditions for asymptotic normality, providing valuable insights into the role of nuisance function estimation when the instrument is continuous. Additionally, we present practical procedures for bandwidth selection and variance estimation. Through extensive simulations, we demonstrate the advantages of the proposed nonparametric estimators. Finally, we apply our methods to data where excess travel time is an instrument for patients' likelihood of receiving care at specialized health care facilities. We use this instrument to estimate the effect of delivering at low-quality neonatal intensive care units (NICUs) on infant mortality.

stat.ME

Learning Smooth Populations of Parameters with Trial Heterogeneity

We consider the classical problem of estimating the mixing distribution of binomial mixtures, but under trial heterogeneity and smoothness. This problem has been studied extensively when the trial parameter is homogeneous, but not under the more general scenario of heterogeneous trials, and only within a low smoothness regime, where the resulting rates are slow. Under the assumption that the density is s-smooth, we derive fast error rates for the kernel density estimator under trial heterogeneity that depend on the harmonic mean of the trials. Importantly, even when reduced to the homogeneous case, our result improves on the state-of-the-art rate of Ye and Bickel (2021). We also study nonparametric estimation of the difference between two densities, which can be smoother than the individual densities, in both i.i.d. and binomial-mixture settings. Our work is motivated by an application in criminal justice: comparing conviction rates of indigent representation in Pennsylvania. We find that the estimated conviction rates for appointed counsel (court-appointed private attorneys) are generally higher than those for public defenders, potentially due to a confounding factor: appointed counsel are more likely to take on severe cases.

math.ST

Discussion of "Causal and counterfactual views of missing data models" by Razieh Nabi, Rohit Bhattacharya, Ilya Shpitser, & James M. Robins

We congratulate Nabi et al. (2022) on their impressive and insightful paper, which illustrates the benefits of using causal/counterfactual perspectives and tools in missing data problems. This paper represents an important approach to missing-not-at-random (MNAR) problems, exploiting nonparametric independence restrictions for identification, as opposed to parametric/semiparametric models, or resorting to sensitivity analysis. Crucially, the authors represent these restrictions with missing data directed acyclic graphs (m-DAGs), which can be useful to determine identification in complex and interesting MNAR models. In this discussion we consider: (i) how/whether other tools from causal inference could be useful in missing data problems, (ii) problems that combine both missing data and causal inference together, and (iii) some work on estimation in one of the authors' example MNAR models.

stat.ME