SearcharxivSearch

arXiv subjects

James Robins

Publications and source records attributed to James Robins.

At least 19 recordsLinked to original sources

Causal Inference: A Tale of Three Frameworks

Causal inference is a central goal across many scientific disciplines. Over the past several decades, three major frameworks have emerged to formalize causal questions and guide their analysis: the potential outcomes framework, structural equation models, and directed acyclic graphs. Although these frameworks differ in language, assumptions, and philosophical orientation, they often lead to compatible or complementary insights. This paper provides a comparative introduction to the three frameworks, clarifying their connections, highlighting their distinct strengths and limitations, and illustrating how they can be used together in practice. The discussion is aimed at researchers and graduate students with some background in statistics or causal inference who are seeking a conceptual foundation for applying causal methods across a range of substantive domains.

math.ST

Structural Nested Mean Models Under Parallel Trends Assumptions

We link and extend two approaches to estimating time-varying treatment effects on repeated continuous outcomes--time-varying Difference in Differences (DiD; see Roth et al. (2023) and Chaisemartin et al. (2023) for reviews) and Structural Nested Mean Models (SNMMs; see Vansteelandt and Joffe (2014) for a review). In particular, we show that SNMMs, previously known to be nonparametrically identified under a no unobserved confounding assumption, are also identified under a conditional parallel trends assumption similar to those typically used to justify time-varying DiD methods (but more amenable to time-varying confounding). Because SNMMs model a broader set of causal estimands, our results allow practitioners of time-varying DiD approaches to address additional types of substantive questions under similar assumptions. SNMMs enable estimation of time-varying effect heterogeneity, lasting effects of a `blip' of treatment at a single time point, effects of sustained interventions (possibly on continuous or multi-dimensional treatments) when treatment repeatedly changes value in the data, controlled direct effects, effects of dynamic treatment strategies that depend on covariate history, and more. We provide a method for sensitivity analysis to violations of our parallel trends assumption. We further explain how to estimate optimal treatment regimes via optimal regime SNMMs under parallel trends assumptions plus an assumption that there is no effect modification by unobserved confounders. Finally, we illustrate our methods with real data applications estimating effects of Medicaid expansion on uninsurance rates, effects of floods on flood insurance take-up, and effects of sustained changes in temperature on crop yields.

stat.ME

Double-robust and efficient methods for estimating the causal effects of a binary treatment

We consider the problem of estimating the effects of a binary treatment on a continuous outcome of interest from observational data in the absence of confounding by unmeasured factors. We provide a new estimator of the population average treatment effect (ATE) based on the difference of novel double-robust (DR) estimators of the treatment-specific outcome means. We compare our new estimator with previously estimators both theoretically and via simulation. DR-difference estimators may have poor finite sample behavior when the estimated propensity scores in the treated and untreated do not overlap. We therefore propose an alternative approach, which can be used even in this unfavorable setting, based on locally efficient double-robust estimation of a semiparametric regression model for the modification on an additive scale of the magnitude of the treatment effect by the baseline covariates $X$. In contrast with existing methods, our approach simultaneously provides estimates of: i) the average treatment effect in the total study population, ii) the average treatment effect in the random subset of the population with overlapping estimated propensity scores, and iii) the treatment effect at each level of the baseline covariates $X$. When the covariate vector $X$ is high dimensional, one cannot be certain, owing to lack of power, that given models for the propensity score and for the regression of the outcome on treatment and $X$ used in constructing our DR estimators are nearly correct, even if they pass standard goodness of fit tests. Therefore to select among candidate models, we propose a novel approach to model selection that leverages the DR-nature of our treatment effect estimator and that outperforms cross-validation in a small simulation study.

stat.ME

Doubly Robust Regression Analysis for Data Fusion

This paper investigates the problem of making inference about a parametric model for the regression of an outcome variable $Y$ on covariates $(V,L)$ when data are fused from two separate sources, one which contains information only on $(V, Y)$ while the other contains information only on covariates. This data fusion setting may be viewed as an extreme form of missing data in which the probability of observing complete data $(V,L,Y)$ on any given subject is zero. We have developed a large class of semiparametric estimators, which includes doubly robust estimators, of the regression coefficients in fused data. The proposed method is DR in that it is consistent and asymptotically normal if, in addition to the model of interest, we correctly specify a model for either the data source process under an ignorability assumption, or the distribution of unobserved covariates. We evaluate the performance of our various estimators via an extensive simulation study, and apply the proposed methods to investigate the relationship between net asset value and total expenditure among U.S. households in 1998, while controlling for potential confounders including income and other demographic variables.

stat.ME

On the multiply robust estimation of the mean of the g-functional

We study multiply robust (MR) estimators of the longitudinal g-computation formula of Robins (1986). In the first part of this paper we review and extend the recently proposed parametric multiply robust estimators of Tchetgen-Tchetgen (2009) and Molina, Rotnitzky, Sued and Robins (2017). In the second part of the paper we derive multiply and doubly robust estimators that use non-parametric machine-learning (ML) estimators of nuisance functions in lieu of parametric models. We use sample splitting to avoid the need for Donsker conditions, thereby allowing an analyst to select the ML algorithms of their choosing. We contrast the asymptotic behavior of our non-parametric doubly robust and multiply robust estimators. In particular, we derive formulas for their asymptotic bias. Examining these formulas we conclude that although, under certain data generating laws, the rate at which the bias of the MR estimator converges to zero can exceed that of the DR estimator, nonetheless, under most laws, the bias of the DR and MR estimators converge to zero at the same rate.

stat.ME

Instrumental variables as bias amplifiers with general outcome and confounding

Drawing causal inference with observational studies is the central pillar of many disciplines. One sufficient condition for identifying the causal effect is that the treatment-outcome relationship is unconfounded conditional on the observed covariates. It is often believed that the more covariates we condition on, the more plausible this unconfoundedness assumption is. This belief has had a huge impact on practical causal inference, suggesting that we should adjust for all pretreatment covariates. However, when there is unmeasured confounding between the treatment and outcome, estimators adjusting for some pretreatment covariate might have greater bias than estimators without adjusting for this covariate. This kind of covariate is called a bias amplifier, and includes instrumental variables that are independent of the confounder, and affect the outcome only through the treatment. Previously, theoretical results for this phenomenon have been established only for linear models. We fill in this gap in the literature by providing a general theory, showing that this phenomenon happens under a wide class of models satisfying certain monotonicity assumptions. We further show that when the treatment follows an additive or multiplicative model conditional on the instrumental variable and the confounder, these monotonicity assumptions can be interpreted as the signs of the arrows of the causal diagrams.

math.ST

Adaptive Estimation of Nonparametric Functionals

We provide general adaptive upper bounds for estimating nonparametric functionals based on second order U-statistics arising from finite dimensional approximation of the infinite dimensional models. We then provide examples of functionals for which the theory produces rate optimally matching adaptive upper and lower bounds. Our results are automatically adaptive in both parametric and nonparametric regimes of estimation and are automatically adaptive and semiparametric efficient in the regime of parametric convergence rate.

math.ST

Double/Debiased Machine Learning for Treatment and Causal Parameters

Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal parameters. Examples of such parameters include individual regression coefficients, average treatment effects, average lifts, and demand or supply elasticities. In fact, estimates of such causal parameters obtained via naively plugging ML estimators into estimating equations for such parameters can behave very poorly due to the regularization bias. Fortunately, this regularization bias can be removed by solving auxiliary prediction problems via ML tools. Specifically, we can form an orthogonal score for the target low-dimensional parameter by combining auxiliary and main ML predictions. The score is then used to build a de-biased estimator of the target parameter which typically will converge at the fastest possible 1/root(n) rate and be approximately unbiased and normal, and from which valid confidence intervals for these parameters of interest may be constructed. The resulting method thus could be called a "double ML" method because it relies on estimating primary and auxiliary predictive models. In order to avoid overfitting, our construction also makes use of the K-fold sample splitting, which we call cross-fitting. This allows us to use a very broad set of ML predictive methods in solving the auxiliary and main prediction problems, such as random forest, lasso, ridge, deep neural nets, boosted trees, as well as various hybrids and aggregators of these methods.

stat.ML

Semiparametric Estimation with Data Missing Not at Random Using an Instrumental Variable

Missing data occur frequently in empirical studies in health and social sciences, often compromising our ability to make accurate inferences. An outcome is said to be missing not at random (MNAR) if, conditional on the observed variables, the missing data mechanism still depends on the unobserved outcome. In such settings, identification is generally not possible without imposing additional assumptions. Identification is sometimes possible, however, if an instrumental variable (IV) is observed for all subjects which satisfies the exclusion restriction that the IV affects the missingness process without directly influencing the outcome. In this paper, we provide necessary and sufficient conditions for nonparametric identification of the full data distribution under MNAR with the aid of an IV. In addition, we give sufficient identification conditions that are more straightforward to verify in practice. For inference, we focus on estimation of a population outcome mean, for which we develop a suite of semiparametric estimators that extend methods previously developed for data missing at random. Specifically, we propose inverse probability weighted estimation, outcome regression-based estimation and doubly robust estimation of the mean of an outcome subject to MNAR. For illustration, the methods are used to account for selection bias induced by HIV testing refusal in the evaluation of HIV seroprevalence in Mochudi, Botswana, using interviewer characteristics such as gender, age and years of experience as IVs.

stat.ME

Technical Report: Higher Order Influence Functions and Minimax Estimation of Nonlinear Functionals

Robins et al, 2008, published a theory of higher order influence functions for inference in semi- and non-parametric models. This paper is a comprehensive manuscript from which Robins et al, was drawn. The current paper includes many results and proofs that were not included in Robins et al due to space limitation. Particular results contained in the present paper that were not reported in Robins et al include the following. Given a set of functionals and their corresponding higher order influence functions, we show how to derive the higher order influence function of their product. We apply this result to obtain higher order influence functions and associated estimators for the mean of a response Y subject to monotone missingness under missing at random. These results also apply to estimating the causal effect of a time dependent treatment on an outcome Y in the presence of time-varying confounding. Finally, we include an appendix that contains proofs for all theorems that were stated without proof in Robins et al, 2008. The initial part of the paper is closely related to Robins et al, the latter parts differ.

stat.ME

Higher Order Estimating Equations for High-dimensional Models

We introduce a new method of estimation of parameters in semiparametric and nonparametric models. The method is based on estimating equations that are $U$-statistics in the observations. The $U$-statistics are based on higher order influence functions that extend ordinary linear influence functions of the parameter of interest, and represent higher derivatives of this parameter. For parameters for which the representation cannot be perfect the method leads to a bias-variance trade-off, and results in estimators that converge at a slower than $\sqrt n$-rate. In a number of examples the resulting rate can be shown to be optimal. We are particularly interested in estimating parameters in models with a nuisance parameter of high dimension or low regularity, where the parameter of interest cannot be estimated at $\sqrt n$-rate, but we also consider efficient $\sqrt n$-estimation using novel nonlinear estimators. The general approach is applied in detail to the example of estimating a mean response when the response is not always observed.

stat.ME

Asymptotic Normality of Quadratic Estimators

We prove conditional asymptotic normality of a class of quadratic U-statistics that are dominated by their degenerate second order part and have kernels that change with the number of observations. These statistics arise in the construction of estimators in high-dimensional semi- and non-parametric models, and in the construction of nonparametric confidence sets. This is illustrated by estimation of the integral of a square of a density or regression function, and estimation of the mean response with missing data. We show that estimators are asymptotically normal even in the case that the rate is slower than the square root of the observations.

stat.ME

Lepski's Method and Adaptive Estimation of Nonlinear Integral Functionals of Density

We study the adaptive minimax estimation of non-linear integral functionals of a density and extend the results obtained for linear and quadratic functionals to general functionals. The typical rate optimal non-adaptive minimax estimators of "smooth" non-linear functionals are higher order U-statistics. Since Lepski's method requires tight control of tails of such estimators, we bypass such calculations by a modification of Lepski's method which is applicable in such situations. As a necessary ingredient, we also provide a method to control higher order moments of minimax estimator of cubic integral functionals. Following a standard constrained risk inequality method, we also show the optimality of our adaptation rates.

math.ST

Identification and Inference for Marginal Average Treatment Effect on the Treated With an Instrumental Variable

In observational studies, treatments are typically not randomized and therefore estimated treatment effects may be subject to confounding bias. The instrumental variable (IV) design plays the role of a quasi-experimental handle since the IV is associated with the treatment and only affects the outcome through the treatment. In this paper, we present a novel framework for identification and inference using an IV for the marginal average treatment effect amongst the treated (ETT) in the presence of unmeasured confounding. For inference, we propose three different semiparametric approaches: (i) inverse probability weighting (IPW), (ii) outcome regression (OR), and (iii) doubly robust (DR) estimation, which is consistent if either (i) or (ii) is consistent, but not necessarily both. A closed-form locally semiparametric efficient estimator is obtained in the simple case of binary IV and outcome and the efficiency bound is derived for the more general case.

stat.ME

Efficient Nonparametric Conformal Prediction Regions

We investigate and extend the conformal prediction method due to Vovk,Gammerman and Shafer (2005) to construct nonparametric prediction regions. These regions have guaranteed distribution free, finite sample coverage, without any assumptions on the distribution or the bandwidth. Explicit convergence rates of the loss function are established for such regions under standard regularity conditions. Approximations for simplifying implementation and data driven bandwidth selection methods are also discussed. The theoretical properties of our method are demonstrated through simulations.

math.ST

Higher order influence functions and minimax estimation of nonlinear functionals

We present a theory of point and interval estimation for nonlinear functionals in parametric, semi-, and non-parametric models based on higher order influence functions (Robins (2004), Section 9; Li et al. (2004), Tchetgen et al. (2006), Robins et al. (2007)). Higher order influence functions are higher order U-statistics. Our theory extends the first order semiparametric theory of Bickel et al. (1993) and van der Vaart (1991) by incorporating the theory of higher order scores considered by Pfanzagl (1990), Small and McLeish (1994) and Lindsay and Waterman (1996). The theory reproduces many previous results, produces new non-$\sqrt{n}$ results, and opens up the ability to perform optimal non-$\sqrt{n}$ inference in complex high dimensional models. We present novel rate-optimal point and interval estimators for various functionals of central importance to biostatistics in settings in which estimation at the expected $\sqrt{n}$ rate is not possible, owing to the curse of dimensionality. We also show that our higher order influence functions have a multi-robustness property that extends the double robustness property of first order influence functions described by Robins and Rotnitzky (2001) and van der Laan and Robins (2003).

math.ST

Causal Models for Estimating the Effects of Weight Gain on Mortality

Suppose, contrary to fact, in 1950, we had put the cohort of 18 year old non-smoking American men on a stringent mandatory diet that guaranteed that no one would ever weigh more than their baseline weight established at age 18. How would the counter-factual mortality of these 18 year olds have compared to their actual observed mortality through 2007? We describe in detail how this counterfactual contrast could be estimated from longitudinal epidemiologic data similiar to that stored in the electronic medical records of a large HMO by applying g-estimation to a novel structural nested model. Our analytic approach differs from any alternative approach in that in that, in the abscence of model misspecification, it can successfully adjust for (i) measured time-varying confounders such as exercise, hypertension and diabetes that are simultaneously intermediate variables on the causal pathway from weight gain to death and determinants of future weight gain, (ii) unmeasured confounding by undiagnosed preclinical disease (i.e reverse causation) that can cause both poor weight gain and premature mortality [provided an upper bound can be specified for the maximum length of time a subject may suffer from a subclinical illness severe enough to affect his weight without the illness becomes clinically manifest], and (iii) the prescence of particular identifiable subgroups, such as those suffering from serious renal, liver, pulmonary, and/or cardiac disease, in whom confounding by unmeasured prognostic factors so severe as to render useless any attempt at direct analytic adjustment.

stat.AP