SearcharxivSearch

arXiv subjects

Andrew Ying

Publications and source records attributed to Andrew Ying.

At least 19 recordsLinked to original sources

An Instrumental Variable Approach to Account for Informative Treatment Switching in Real-world Evidence

Reproducible and generalizable assessment of treatment decisions requires principled handling of subsequent treatment switching that may inform expected outcomes and shift across cohorts and over time. To effectively account for informative treatment switching, we propose an instrumental variable approach that characterizes the poorly documented expected outcomes at switching as unmeasured confounding. After establishing the baseline treatment as a viable instrumental variable, we constructed an estimating equation based on the association between the centered instrumental variable and a martingale style residual process that identifies the treatment effect under structural cumulative survival model. Our proposed method is doubly robust, i.e., valid whenever either of baseline propensity model or no-switching outcome model is consistently estimated. A co-training of treatment effect parameter and survival outcome regression model eliminated the requirement of observing a no-switching subset under semi-parametric additive hazards models. We further developed an baseline-survival-corrected cross-fitting approach to incorporate general machine learning models for estimating nuisance models. Numerical results demonstrated the validity of our method in various settings when a basket of benchmark solutions produced biased or contradictory results. We applied our method to comparison of high-efficacy vs standard efficacy disease modifying treatments as the second line therapy of multiple sclerosis.

stat.ME

Efficient Inference for Incremental Causal Effects of Time to Treatment

We consider continuous time to treatment initiation. This can commonly occur in preventive medicine, such as disease screening and vaccination; it can also occur with non-fatal health conditions such as HIV infection without the onset of AIDS. While traditional causal inference focused on `when to treat' and its effects, we consider the incremental causal effect when the intensity of time to treatment initiation is intervened upon. We derive the efficient influence function for this estimand and develop an estimation framework that accommodates flexible machine learning methods while achieving fast convergence rates. Valid confidence bands are obtained leveraging empirical process theory. We illustrate our approach via simulation, and apply it to cervical cancer screening data to study the incremental effect of time to subsequent HPV testing on cervical intraepithelial neoplasia detection.

stat.ME

Proximal Survival Analysis for Dependent Left Truncation

In prevalent cohort studies with delayed entry, time-to-event outcomes are often subject to left truncation where only subjects that have not experienced the event at study entry are included, leading to selection bias. Existing methods for handling left truncation mostly rely on the (quasi-)independence assumption or the weaker conditional (quasi-)independence assumption which assumes that conditional on observed covariates, the left truncation time and the event time are independent on the observed region. In practice, however, our analysis of the Honolulu Asia Aging Study (HAAS) suggests that the conditional quasi-independence assumption may fail because measured covariates often serve only as imperfect proxies for the underlying mechanisms, such as latent health status, that induce dependence between truncation and event times. To address this gap, we propose a proximal weighting identification framework that admits the dependence-inducing factors may not be fully observed. We then construct an estimator based on the framework and study its asymptotic properties. We examine the finite sample performance of the proposed estimator by comprehensive simulations, and apply it to analyzing the cognitive impairment-free survival probabilities using data from the Honolulu Asia Aging Study.

stat.ME

A Liberating Framework from Truncation and Censoring, with Application to Learning Treatment Effects

Time-to-event outcomes are often subject to left truncation and right censoring. While many survival analysis methods have been developed to handle truncation and censoring, majority of the past works require strong independence assumptions. We relax these stringent assumptions through leveraging covariate information together with orthogonal learning, and develop a liberating framework from left truncation and right censoring so that desirable properties like double robustness can be immediately transferred from settings without truncation or censoring. To illustrate its generality and ease to use, the framework is applied to estimation of the average treatment effect (ATE) and the conditional average treatment effect (CATE). For the ATE, we establish both model and rate double robustness under confounding, truncation and censoring; for the CATE, we show that the orthogonal and the doubly robust learners under these three sources of bias can achieve oracle rate of convergence. We study the estimators both theoretically and through extensive simulation, and apply them to analyzing the effect of mid-life heavy drinking on late life cognitive impairment free survival, using data from the Honolulu Asia Aging Study.

stat.ME

Incremental Causal Effect for Time to Treatment Initialization

We consider time to treatment initialization. This can commonly occur in preventive medicine, such as disease screening and vaccination; it can also occur with non-fatal health conditions such as HIV infection without the onset of AIDS; or in tech industry where items wait to be reviewed manually as abusive or not, etc. While traditional causal inference focused on `when to treat' and its effects, including their possible dependence on subject characteristics, we consider the incremental causal effect when the intensity of time to treatment initialization is intervened upon. We provide identification of the incremental causal effect without the commonly required positivity assumption, as well as an estimation framework using inverse probability weighting. We illustrate our approach via simulation, and apply it to a rheumatoid arthritis study to evaluate the incremental effect of time to start methotrexate on joint pain.

stat.ME

Dynamic Treatment Effects under Functional Longitudinal Studies

Establishing causality is a fundamental goal in fields like medicine and social sciences. While randomized controlled trials are the gold standard for causal inference, they are not always feasible or ethical. Observational studies can serve as alternatives but introduce confounding biases, particularly in complex longitudinal data, where treatment-confounder feedback complicates analysis. The challenge increases with Dynamic Treatment Regimes (DTRs), where treatment allocation depends on rich historical patient data. The advent of real-time healthcare monitoring technologies, such as MIMIC-IV and Continuous Glucose Monitoring (CGM), has popularized Functional Longitudinal Data (FLD). However, there is yet no investigate of causal inference for FLD with DTRs. In this paper, we address it by developing a population-level framework for functional longitudinal data, accommodating DTRs. To that end, we define the potential outcomes and causal effects of interest. We then develop identification assumptions, and derive g-computation, inverse probability weighting, and doubly robust formulas through novel applications of stochastic process and measure theory. We further show that our framework is nonparametric and compute the efficient influence curve using semiparametric theory. Last, we illustrate our framework's potential through Monte Carlo simulations.

math.ST

Deepening the Understanding of Double Robustness Geometrically

Double robustness (DR) is a widely-used property of estimators that provides protection against model misspecification and slow convergence of nuisance functions. Despite its widespread application, the theoretical foundation of DR remains underexplored. While DR is a property of global invariance along both nuisance directions, it is often implied by influence curves (ICs), which only have zero first-order derivatives in those directions locally. On the other hand, some literature proved the absence of DR estimating functions for the same estimand, under one parameterization yet was able to find one under another parameterization, highlighting the nuances in parameterization. In this short communication, we address two key questions: (1) Why do ICs frequently imply DR ``for free''? (2) Under what conditions would a given statistical model and parameterization support or prevent the existence of DR estimators? Using tools from semiparametric theory, we show that convexity is the crucial property that enables influence curves to imply DR. We then derive necessary and sufficient conditions for the existence of DR estimators. Our main contribution also lies in the novel geometric interpretation of DR using information geometry, a discipline devoted to integrating global differential geometry with statistical analysis. By leveraging concepts such as parallel transport, m-flatness, and m-curvature freeness, we characterize DR in terms of invariance along submanifolds. This geometric perspective deepens the understanding of when and why DR estimators exist.

math.ST

On Defense of the Hazard Ratio

In this short communication, we describe the recent debate on whether the hazard function should be used for causal inference in time-to-event studies and consider three different potential outcomes frameworks (by Rubin, Robins, and Pearl, respectively) as well as use the single-world intervention graph to show mathematically that the hazard function has causal interpretations under all three frameworks. In addition, we argue that the hazard ratio over time can provide a useful interpretation in practical settings.

math.ST

Asymptotic Theory for Doubly Robust Estimators with Continuous-Time Nuisance Parameters

Doubly robust estimators have gained widespread popularity in various fields due to their ability to provide unbiased estimates under model misspecification. However, the asymptotic theory for doubly robust estimators with continuous-time nuisance parameters remains largely unexplored. In this short communication, we address this gap by developing a general asymptotic theory for a class of doubly robust estimating equations involving stochastic processes and Riemann-Stieltjes integrals. We introduce generic assumptions on the nuisance parameter estimators that ensure the consistency and asymptotic normality of the resulting doubly robust estimator. Our results cover both the model doubly robust estimator, which relies on parametric or semiparametric models, and the rate doubly robust estimator, which allows for flexible machine learning methods. We discuss the implications of our findings and highlight the key differences between the continuous-time setting and the classical theory for doubly robust estimators. Our work provides a solid theoretical foundation for the use of doubly robust estimators in complex settings with continuous-time nuisance parameters, paving the way for future research and applications.

math.ST

Proximal Survival Analysis to Handle Dependent Right Censoring

Many epidemiological and clinical studies aim at analyzing a time-to-event endpoint. A common complication is right censoring. In some cases, it arises because subjects are still surviving after the study terminates or move out of the study area, in which case right censoring is typically treated as independent or non-informative. Such an assumption can be further relaxed to conditional independent censoring by leveraging possibly time-varying covariate information, if available, assuming censoring and failure time are independent among covariate strata. In yet other instances, events may be censored by other competing events like death and are associated with censoring possibly through prognoses. Realistically, measured covariates can rarely capture all such associations with certainty. For such dependent censoring, often covariate measurements are at best proxies of underlying prognoses. In this paper, we establish a nonparametric identification framework by formally admitting that conditional independent censoring may fail in practice and accounting for covariate measurements as imperfect proxies of underlying association. The framework suggests adaptive estimators which we give generic assumptions under which they are consistent, asymptotically normal, and doubly robust. We illustrate our framework with concrete settings, where we examine the finite-sample performance of our proposed estimators via a Monte-Carlo simulation and apply them to the SEER-Medicare dataset.

stat.ME

Doubly Robust Estimation under Covariate-Induced Dependent Left Truncation

In prevalent cohort studies with follow-up, the time-to-event outcome is subject to left truncation leading to selection bias. For estimation of the distribution of time-to-event, conventional methods adjusting for left truncation tend to rely on the (quasi-)independence assumption that the truncation time and the event time are "independent" on the observed region. This assumption is violated when there is dependence between the truncation time and the event time possibly induced by measured covariates. Inverse probability of truncation weighting leveraging covariate information can be used in this case, but it is sensitive to misspecification of the truncation model. In this work, we apply the semiparametric theory to find the efficient influence curve of an expected (arbitrarily transformed) survival time in the presence of covariate-induced dependent left truncation. We then use it to construct estimators that are shown to enjoy double-robustness properties. Our work represents the first attempt to construct doubly robust estimators in the presence of left truncation, which does not fall under the established framework of coarsened data where doubly robust approaches are developed. We provide technical conditions for the asymptotic properties that appear to not have been carefully examined in the literature for time-to-event data, and study the estimators via extensive simulation. We apply the estimators to two data sets from practice, with different right-censoring patterns.

stat.ME

Causal Identification for Complex Functional Longitudinal Studies

Real-time monitoring in modern medical research introduces functional longitudinal data, characterized by continuous-time measurements of outcomes, treatments, and confounders. This complexity leads to uncountably infinite treatment-confounder feedbacks, which traditional causal inference methodologies cannot handle. Inspired by the coarsened data framework, we adopt stochastic process theory, measure theory, and net convergence to propose a nonparametric causal identification framework. This framework generalizes classical g-computation, inverse probability weighting, and doubly robust formulas, accommodating time-varying outcomes subject to mortality and censoring for functional longitudinal data. We examine our framework through Monte Carlo simulations. Our approach addresses significant gaps in current methodologies, providing a solution for functional longitudinal data and paving the way for future estimation work in this domain.

stat.ME

Proximal Causal Inference for Marginal Counterfactual Survival Curves

Contrasting marginal counterfactual survival curves across treatment arms is an effective and popular approach for inferring the causal effect of an intervention on a right-censored time-to-event outcome. A key challenge to drawing such inferences in observational settings is the possible existence of unmeasured confounding, which may invalidate most commonly used methods that assume no hidden confounding bias. In this paper, rather than making the standard no unmeasured confounding assumption, we extend the recently proposed proximal causal inference framework of Miao et al. (2018), Tchetgen et al. (2020), Cui et al. (2020) to obtain nonparametric identification of a causal survival contrast by leveraging observed covariates as imperfect proxies of unmeasured confounders. Specifically, we develop a proximal inverse probability-weighted (PIPW) estimator, the proximal analog of standard IPW, which allows the observed data distribution for the time-to-event outcome to remain completely unrestricted. PIPW estimation relies on a parametric model for a so-called treatment confounding bridge function relating the treatment process to confounding proxies. As a result, PIPW might be sensitive to model misspecification. To improve robustness and efficiency, we also propose a proximal doubly robust estimator and establish uniform consistency and asymptotic normality of both estimators. We conduct extensive simulations to examine the finite sample performance of our estimators, and proposed methods are applied to a study evaluating the effectiveness of right heart catheterization in the intensive care unit of critically ill patients.

stat.ME

A Robust Instrumental Variable Method Accounting for Treatment Switching in Open-Label Randomized Controlled Trials

In a randomized controlled trial, treatment switching (also called contamination or crossover) occurs when a patient initially assigned to one treatment arm changes to another arm during the course of follow-up. Overlooking treatment switching might substantially bias the evaluation of treatment efficacy or safety. To account for treatment switching, instrumental variable (IV) methods by leveraging the initial randomized assignment as an IV serve as natural adjustment methods because they allow dependent treatment switching possibly due to underlying prognoses. However, the ``exclusion restriction'' assumption for IV methods, which requires the initial randomization to have no direct effect on the outcome, remains questionable, especially for open-label trials. We propose a robust instrumental variable estimator circumventing such a caveat. We derive large-sample properties of our proposed estimator, along with inferential tools. We conduct extensive simulations to examine the finite performance of our estimator and its associated inferential tools. An R package ``ivsacim'' implementing all proposed methods is freely available on R CRAN. We apply the estimator to evaluate the treatment effect of Nucleoside Reverse Transcriptase Inhibitors (NRTIs) on a safety outcome in the Optimized Treatment That Includes or Omits NRTIs trial.

stat.ME

Marginal Structural Illness-Death Models for Semi-Competing Risks Data

The three state illness death model has been established as a general approach for regression analysis of semi competing risks data. For observational data the marginal structural models (MSM) are a useful tool, under the potential outcomes framework to define and estimate parameters with causal interpretations. In this paper we introduce a class of marginal structural illness death models for the analysis of observational semi competing risks data. We consider two specific such models, the Markov illness death MSM and the frailty based Markov illness death MSM. For interpretation purposes, risk contrasts under the MSMs are defined. Inference under the illness death MSM can be carried out using estimating equations with inverse probability weighting, while inference under the frailty based illness death MSM requires a weighted EM algorithm. We study the inference procedures under both MSMs using extensive simulations, and apply them to the analysis of mid life alcohol exposure on late life cognitive impairment as well as mortality using the Honolulu Asia Aging Study data set. The R codes developed in this work have been implemented in the R package semicmprskcoxmsm that is publicly available on CRAN.

stat.ME

Proximal Causal Inference for Complex Longitudinal Studies

A standard assumption for causal inference about the joint effects of time-varying treatment is that one has measured sufficient covariates to ensure that within covariate strata, subjects are exchangeable across observed treatment values, also known as "sequential randomization assumption (SRA)". SRA is often criticized as it requires one to accurately measure all confounders. Realistically, measured covariates can rarely capture all confounders with certainty. Often covariate measurements are at best proxies of confounders, thus invalidating inferences under SRA. In this paper, we extend the proximal causal inference (PCI) framework of Miao et al. (2018) to the longitudinal setting under a semiparametric marginal structural mean model (MSMM). PCI offers an opportunity to learn about joint causal effects in settings where SRA based on measured time-varying covariates fails, by formally accounting for the covariate measurements as imperfect proxies of underlying confounding mechanisms. We establish nonparametric identification with a pair of time-varying proxies and provide a corresponding characterization of regular and asymptotically linear estimators of the parameter indexing the MSMM, including a rich class of doubly robust estimators, and establish the corresponding semiparametric efficiency bound for the MSMM. Extensive simulation studies and a data application illustrate the finite sample behavior of proposed methods.

stat.ME

Minimax Kernel Machine Learning for a Class of Doubly Robust Functionals with Application to Proximal Causal Inference

Robins et al. (2008) introduced a class of influence functions (IFs) which could be used to obtain doubly robust moment functions for the corresponding parameters. However, that class does not include the IF of parameters for which the nuisance functions are solutions to integral equations. Such parameters are particularly important in the field of causal inference, specifically in the recently proposed proximal causal inference framework of Tchetgen Tchetgen et al. (2020), which allows for estimating the causal effect in the presence of latent confounders. In this paper, we first extend the class of Robins et al. to include doubly robust IFs in which the nuisance functions are solutions to integral equations. Then we demonstrate that the double robustness property of these IFs can be leveraged to construct estimating equations for the nuisance functions, which enables us to solve the integral equations without resorting to parametric models. We frame the estimation of the nuisance functions as a minimax optimization problem. We provide convergence rates for the nuisance functions and conditions required for asymptotic linearity of the estimator of the parameter of interest. The experiment results demonstrate that our proposed methodology leads to robust and high-performance estimators for average causal effect in the proximal causal inference framework.

stat.ML

A New Causal Approach to Account for Treatment Switching in Randomized Experiments under a Structural Cumulative Survival Model

Treatment switching in a randomized controlled trial is said to occur when a patient randomized to one treatment arm switches to another treatment arm during follow-up. This can occur at the point of disease progression, whereby patients in the control arm may be offered the experimental treatment. It is widely known that failure to account for treatment switching can seriously dilute the estimated effect of treatment on overall survival. In this paper, we aim to account for the potential impact of treatment switching in a re-analysis evaluating the treatment effect of NucleosideReverse Transcriptase Inhibitors (NRTIs) on a safety outcome (time to first severe or worse sign or symptom) in participants receiving a new antiretroviral regimen that either included or omitted NRTIs in the Optimized Treatment That Includes or OmitsNRTIs (OPTIONS) trial. We propose an estimator of a treatment causal effect under a structural cumulative survival model (SCSM) that leverages randomization as an instrumental variable to account for selective treatment switching. Unlike Robins' accelerated failure time model often used to address treatment switching, the proposed approach avoids the need for artificial censoring for estimation. We establish that the proposed estimator is uniformly consistent and asymptotically Gaussian under standard regularity conditions. A consistent variance estimator is also given and a simple resampling approach provides uniform confidence bands for the causal difference comparing treatment groups overtime on the cumulative intensity scale. We develop an R package named "ivsacim" implementing all proposed methods, freely available to download from R CRAN. We examine the finite performance of the estimator via extensive simulations.

stat.ME