SearcharxivSearch

arXiv subjects

Shanshan Luo

Publications and source records attributed to Shanshan Luo.

At least 19 recordsLinked to original sources

Selecting among Missingness Models for Sequential Outcomes with Nonignorable Nonresponse

Sequential outcomes in longitudinal studies and multi-wave surveys may be missing not at random at both earlier and later occasions. We study graphical models in which at least one outcome is self-censoring and the response indicator for a later outcome may depend on either the earlier response indicator or the realized earlier outcome. These restrictions define two candidate families under which the relevant full-data distributions are identifiable and whose observed-data models overlap; graphs containing both dependencies form a broader class outside the prespecified comparison. For each candidate family, we establish identification of the full-data distribution under rank or completeness conditions and develop likelihood-based estimation. We then propose a two-stage Vuong-type procedure. The first stage determines whether the candidate models are observationally distinguishable; only after distinguishability is established does the second stage compare their Kullback--Leibler divergences from the true observed-data distribution. We also show that ordinary Wald inference remains asymptotically valid for the selected model-specific functional when the selected model has a fixed positive expected log-likelihood advantage. Simulations evaluate the two-stage procedure across graph classes. We finally apply the procedure to compare candidate missingness models in the Job Corps data and perform downstream functional estimation under the selected model.

stat.ME

Mediation Analysis with Multiple Mediators Subject to Missing Not at Random

Causal mediation analysis serves as a key tool for uncovering the mediating mechanisms linking treatments to outcomes. Existing methods for mediation analysis with multiple mediators typically assume complete observations or missing-at-random and may yield biased estimation when mediator values are missing not at random (MNAR). This paper studies the identification and estimation of causal mediation effects with multiple mediators subject to MNAR missingness. We consider a broad class of MNAR mechanisms in which missingness may depend on unobserved mediators, treatment, covariates, and outcomes. Under a series of increasingly general MNAR mechanisms, we establish identified natural direct and indirect effects, effectively generalizing existing mediation analysis to handle nonignorable missing mediators. Based on the proposed identification framework, we develop estimation procedures for causal mediation effects and evaluate their finite-sample performance through simulation studies. The results demonstrate satisfactory performance across a range of missingness scenarios. An application to data from the National Health and Nutrition Examination Survey(NHANES) illustrates the practical utility of the proposed methodology for investigating mediation pathways in the presence of nonignorable missing data.

stat.ME

Apportioning Causal Responsibility of Two Risk Factors for an Adverse Outcome via Counterfactual Attribution

Unlike traditional causal inference, which prospectively evaluates the effects of causes, apportioning causal responsibility requires a retrospective assessment to deduce the causes of an outcome that has already occurred. This paper proposes a quantitative framework for apportioning causal responsibility between two binary risk factors that jointly contribute to a realized adverse outcome. Ideally, knowing the individual's latent causal type, defined by the potential outcomes under all possible exposure combinations, would allow precise apportionment; however, these potential outcomes cannot be simultaneously observed. We therefore define the average causal responsibility of each risk factor as its expected responsibility over the distribution of latent causal types. Under the assumptions of no confounding and monotonicity, we establish nonparametric identification of this metric when the type-specific responsibilities satisfy a structural balance condition, and derive sharp bounds otherwise. We illustrate the proposed framework using the classic example of lung cancer attributable to smoking and asbestos exposures.

stat.ME

Bidirectional causal inference for binary outcomes in the presence of unmeasured confounding

Bidirectional causal relationships arising from mutual interactions between variables are commonly observed within biomedical, econometrical, and social science contexts. When such relationships are further complicated by unobserved factors, identifying causal effects in both directions becomes especially challenging. For continuous variables, methods that utilize two instrumental variables from both directions have been proposed to explore bidirectional causal effects in linear models. However, the existing techniques are not applicable when the key variables of interest are binary. To address these issues, we propose a structural equation modeling approach that links observed binary variables to continuous latent variables through a constrained mapping. We further establish identification results for bidirectional causal effects using a pair of instrumental variables. Additionally, we develop an estimation method for the corresponding causal parameters. We also conduct sensitivity analysis under scenarios where certain identification conditions are violated. Finally, we apply our approach to investigate the bidirectional causal relationship between heart disease and diabetes, demonstrating its practical utility in biomedical research.

stat.ME

Assessing Interactive Causes of an Occurred Outcome Due to Two Binary Exposures

In contrast to evaluating treatment effects, causal attribution analysis focuses on identifying the key factors responsible for an observed outcome. For two binary exposure variables and a binary outcome variable, researchers need to assess not only the likelihood that an observed outcome was caused by a particular exposure, but also the likelihood that it resulted from the interaction between the two exposures. For example, in the case of a male worker who smoked, was exposed to asbestos, and developed lung cancer, researchers aim to explore whether the cancer resulted from smoking, asbestos exposure, or their interaction. Even in randomized controlled trials, widely regarded as the gold standard for causal inference, identifying and evaluating retrospective causal interactions between two exposures remains challenging. In this paper, we define posterior probabilities to characterize the interactive causes of an observed outcome. We establish the identifiability of posterior probabilities by using a secondary outcome variable that may appear after the primary outcome. We apply the proposed method to the classic case of smoking and asbestos exposure. Our results indicate that for lung cancer patients who smoked and were exposed to asbestos, the disease is primarily attributable to the synergistic effect between smoking and asbestos exposure.

stat.AP

Pseudo-strata learning via maximizing misclassification reward

Online advertising aims to increase user engagement and maximize revenue, but users respond heterogeneously to ad exposure. Some users purchase only when exposed to ads, while others purchase regardless of exposure, and still others never purchase. This heterogeneity can be characterized by latent response types, commonly referred to as principal strata, defined by users' joint potential outcomes under exposure and non-exposure. However, users' true strata are unobserved, making direct analysis infeasible. In this article, instead of learning the true strata, we propose a novel approach that learns users' pseudo-strata by leveraging information from an outcome (revenue) observed after the response (purchase). We construct pseudo-strata to classify users and introduce misclassification rewards to quantify the expected revenue gain of pseudo-strata-based policies relative to true strata. Within a Bayesian classification framework, we learn the pseudo-strata by optimizing the expected revenue. To implement these procedures, we introduce identification assumptions and estimation methods, and establish their large-sample properties. Simulation studies show that the proposed method achieves more accurate strata classification and substantially higher revenue than baselines. We further illustrate the method using a large-scale industrial dataset from the Criteo Predictive Search Platform.

stat.ME

Covariate Balancing Value Estimation for Optimal Individualized Treatment Rules

Learning an optimal individualized treatment rule depends on reliable value comparisons across the candidate class. Standard doubly robust estimators are consistent when either the propensity score or outcome regression model is correctly specified, but they do not directly control the remaining bias in value estimation when both models are misspecified. In this paper, we propose a covariate balancing doubly robust estimator that combines propensity score estimation based on covariate balancing with an outcome regression component selected using an empirical variance criterion based on the influence function. Using prespecified covariate functions, the balancing procedure induces an effective balancing space to which the weighted propensity score error is orthogonal. Consequently, the proposed value estimator is consistent if either the propensity score model is correct or the rule-relevant outcome regression error lies in this space. The latter condition does not require a correctly specified outcome regression model and can hold even when both working models are misspecified. Under correct propensity score specification, the estimator has the smallest asymptotic variance within the proposed covariate balancing doubly robust class. Simulations evaluate value estimation in finite samples and the regret of learned rules, and an application to a leukemia dataset illustrates the proposed method in practice.

stat.ME

A regression-based approach for bidirectional proximal causal inference in the presence of unmeasured confounding

Proxy variables are commonly used in causal inference when unmeasured confounding exists. While most existing proximal methods assume a unidirectional causal relationship between two primary variables, many social and biological systems exhibit complex feedback mechanisms that imply bidirectional causality. In this paper, using regression-based models, we extend the proximal framework to identify bidirectional causal effects in the presence of unmeasured confounding. We establish the identification of bidirectional causal effects and develop a sensitivity analysis method for violations of the proxy structural conditions. Building on this identification result, we derive bidirectional two-stage least squares estimators that are consistent and asymptotically normal under standard regularity conditions. Simulation studies demonstrate that our approach delivers unbiased causal effect estimates and outperforms some standard methods. The simulation results also confirm the reliability of the sensitivity analysis procedure. Applying our methodology to a state-level panel dataset from 1985 to 2014 in the United States, we examine the bidirectional causal effects between abortion rates and murder rates. The analysis reveals a consistent negative effect of abortion rates on murder rates, while also detecting a potential reciprocal effect from murder rates to abortion rates that conventional unidirectional analyses have not considered.

stat.ME

Safe Individualized Treatment Rules with Controllable Harm Rates

Estimating individualized treatment rules (ITRs) is crucial for tailoring interventions in precision medicine. Typical ITR estimation methods rely on conditional average treatment effects (CATEs) to guide treatment assignments. However, such methods overlook individual-level harm within covariate-specific subpopulations, potentially leading many individuals to experience worse outcomes under CATE-based ITRs. In this article, we aim to estimate ITRs that maximize the reward while ensuring that the harm rate induced by the ITR remains below a pre-specified threshold. We first derive the explicit form of the oracle ITR. However, the oracle ITR is not achievable without strong assumptions, as the harm rate is generally unidentifiable due to its dependence on the joint distribution of potential outcomes. To address this, we propose two strategies for estimating ITRs with a harm rate constraint under partial identification and establish their large-sample properties. By accounting for both reward and harm, our method provides a reliable solution for developing ITRs in high-stakes domains where harm is a critical consideration. Extensive simulations demonstrate the effectiveness of the proposed methods in controlling harm rates. We apply the proposed method to analyze two real-world datasets from a new perspective, assessing the potential reduction in harm rate compared with historical interventions.

stat.ME

A generalized tetrad constraint for testing conditional independence given a latent variable

The tetrad constraint is widely used to test whether four observed variables are conditionally independent given a latent variable, based on the fact that if four observed variables following a linear model are mutually independent after conditioning on an unobserved variable, then products of covariances of any two different pairs of these four variables are equal. It is an important tool for discovering a latent common cause or distinguishing between alternative linear causal structures. However, the classical tetrad constraint fails in nonlinear models because the covariance of observed variables cannot capture nonlinear association. In this paper, we propose a generalized tetrad constraint, which establishes a testable implication for conditional independence given a latent variable in nonlinear and nonparametric models. In linear models, this constraint implies the classical tetrad constraint; in nonlinear models, it remains a necessary condition for conditional independence but the classical tetrad constraint no longer is. Based on this constraint, we further propose a formal test, which can control type I error and has power approaching unity under certain conditions. We illustrate the proposed approach via simulations and two real data applications on mental ability tests and on moral attitudes towards dishonesty.

stat.ME

Identification and estimation of causal peer effects using instrumental variables

In social science researches, causal inference regarding peer effects often faces significant challenges due to homophily bias and contextual confounding. For example, unmeasured health conditions (e.g., influenza) and psychological states (e.g., happiness, loneliness) can spread among closely connected individuals, such as couples or siblings. To address these issues, we define four effect estimands for dyadic data to characterize direct effects and spillover effects. We employ dual instrumental variables to achieve nonparametric identification of these causal estimands in the presence of unobserved confounding. We then derive the efficient influence functions for these estimands under the nonparametric model. Additionally, we develop a triply robust and locally efficient estimator that remains consistent even under partial misspecification of the observed data model. The proposed robust estimators can be easily adapted to flexible approaches such as machine learning estimation methods, provided that certain rate conditions are satisfied. Finally, we illustrate our approach through simulations and an empirical application evaluating the peer effects of retirement on fluid cognitive perception among couples.

stat.ME

Causal Inference with Outcomes Truncated by Death and Missing Not at Random

In clinical trials, principal stratification analysis is commonly employed to address the issue of truncation by death, where a subject dies before the outcome can be measured. However, in practice, many survivor outcomes may remain uncollected or be missing not at random, posing a challenge to standard principal stratification analyses. In this paper, we explore the identification, estimation, and bounds of the average treatment effect within a subpopulation of individuals who would potentially survive under both treatment and control conditions. We show that the causal parameter of interest can be identified by introducing a proxy variable that affects the outcome only through the principal strata, while requiring that the treatment variable does not directly affect the missingness mechanism. Subsequently, we propose an approach for estimating causal parameters and derive nonparametric bounds in cases where identification assumptions are violated. We illustrate the performance of the proposed method through simulation studies and a real dataset obtained from a Human Immunodeficiency Virus (HIV) study.

stat.ME

Automating the Selection of Proxy Variables of Unmeasured Confounders

Recently, interest has grown in the use of proxy variables of unobserved confounding for inferring the causal effect in the presence of unmeasured confounders from observational data. One difficulty inhibiting the practical use is finding valid proxy variables of unobserved confounding to a target causal effect of interest. These proxy variables are typically justified by background knowledge. In this paper, we investigate the estimation of causal effects among multiple treatments and a single outcome, all of which are affected by unmeasured confounders, within a linear causal model, without prior knowledge of the validity of proxy variables. To be more specific, we first extend the existing proxy variable estimator, originally addressing a single unmeasured confounder, to accommodate scenarios where multiple unmeasured confounders exist between the treatments and the outcome. Subsequently, we present two different sets of precise identifiability conditions for selecting valid proxy variables of unmeasured confounders, based on the second-order statistics and higher-order statistics of the data, respectively. Moreover, we propose two data-driven methods for the selection of proxy variables and for the unbiased estimation of causal effects. Theoretical analysis demonstrates the correctness of our proposed algorithms. Experimental results on both synthetic and real-world data show the effectiveness of the proposed approach.

cs.LG

Assessing the causes of continuous effects by posterior effects of causes

To evaluate a single cause of a binary effect, Dawid et al. (2014) defined the probability of causation, while Pearl (2015) defined the probabilities of necessity and sufficiency. For assessing the multiple correlated causes of a binary effect, Lu et al. (2023) defined the posterior causal effects based on post-treatment variables. In many scenarios, outcomes are continuous, simply binarizing them and applying previous methods may result in information loss or biased conclusions. To address this limitation, we propose a series of posterior causal estimands for retrospectively evaluating multiple correlated causes from a continuous effect, including posterior intervention effects, posterior total causal effects, and posterior natural direct effects. Under the assumptions of sequential ignorability, monotonicity, and perfect positive rank, we show that the posterior causal estimands of interest are identifiable and present the corresponding identification equations. We also provide a simple but effective estimation procedure and establish the asymptotic properties of the proposed estimators. An artificial hypertension example and a real developmental toxicity dataset are employed to illustrate our method.

stat.ME

Efficiency-improved doubly robust estimation with non-confounding predictive covariates

In observational studies, covariates with substantial missing data are often omitted, despite their strong predictive capabilities. These excluded covariates are generally believed not to simultaneously affect both treatment and outcome, indicating that they are not genuine confounders and do not impact the identification of the average treatment effect (ATE). In this paper, we introduce an alternative doubly robust (DR) estimator that fully leverages non-confounding predictive covariates to enhance efficiency, while also allowing missing values in such covariates. Beyond the double robustness property, our proposed estimator is designed to be more efficient than the standard DR estimator. Specifically, when the propensity score model is correctly specified, it achieves the smallest asymptotic variance among the class of DR estimators, and brings additional efficiency gains by further integrating predictive covariates. Simulation studies demonstrate the notable performance of the proposed estimator over current popular methods. An illustrative example is provided to assess the effectiveness of right heart catheterization (RHC) for critically ill patients.

stat.ME

On the Comparative Analysis of Average Treatment Effects Estimation via Data Combination

There is growing interest in exploring causal effects in target populations via data combination. However, most approaches are tailored to specific settings and lack comprehensive comparative analyses. In this article, we focus on a typical scenario involving a source dataset and a target dataset. We first design six settings under covariate shift and conduct a comparative analysis by deriving the semiparametric efficiency bounds for the ATE in the target population. We then extend this analysis to six new settings that incorporate both covariate shift and posterior drift. Our study uncovers the key factors that influence efficiency gains and the ``effective sample size" when combining two datasets, with a particular emphasis on the roles of the variance ratio of potential outcomes between datasets and the derivatives of the posterior drift function. To the best of our knowledge, this is the first paper that explicitly explores the role of the posterior drift functions in causal inference. Additionally, we also propose novel methods for conducting sensitivity analysis to address violations of transportability between the two datasets. We empirically validate our findings by constructing locally efficient estimators and conducting extensive simulations. We demonstrate the proposed methods in two real-world applications.

stat.ME

Multiply robust estimation of causal effects using linked data

Unmeasured confounding presents a common challenge in observational studies, potentially making standard causal parameters unidentifiable without additional assumptions. Given the increasing availability of diverse data sources, exploiting data linkage offers a potential solution to mitigate unmeasured confounding within a primary study of interest. However, this approach often introduces selection bias, as data linkage is feasible only for a subset of the study population. To address this concern, we explore three nonparametric identification strategies under the assumption that a unit' s inclusion in the linked cohort is determined solely by the observed confounders, while acknowledging that the ignorability assumption may depend on some partially unobserved covariates. The existence of multiple identification strategies motivates the development of estimators that effectively capture distinct components of the observed data distribution. Appropriately combining these estimators yields triply robust estimators for the average treatment effect. These estimators remain consistent if at least one of the three distinct parts of the observed data law is correct. Moreover, they are locally efficient if all the models are correctly specified. We evaluate the proposed estimators using simulation studies and real data analysis.

stat.ME

Identifying Causal Effects Using Instrumental Variables from the Auxiliary Dataset

Instrumental variable approaches have gained popularity for estimating causal effects in the presence of unmeasured confounders. However, the availability of instrumental variables in the primary dataset is often challenged due to stringent and untestable assumptions. This paper presents a novel method to identify and estimate causal effects by utilizing instrumental variables from the auxiliary dataset, incorporating a structural equation model, even in scenarios with nonlinear treatment effects. Our approach involves using two datasets: one called the primary dataset with joint observations of treatment and outcome, and another auxiliary dataset providing information about the instrument and treatment. Our strategy differs from most existing methods by not depending on the simultaneous measurements of instrument and outcome. The central idea for identifying causal effects is to establish a valid substitute through the auxiliary dataset, addressing unmeasured confounders. This is achieved by developing a control function and projecting it onto the function space spanned by the treatment variable. We then propose a three-step estimator for estimating causal effects and derive its asymptotic results. We illustrate the proposed estimator through simulation studies, and the results demonstrate favorable performance. We also conduct a real data analysis to evaluate the causal effect between vitamin D status and body mass index.

stat.ME