Searcharxiv⌕ Search

arXiv subjects

S. Ghazaleh Dashti

Publications and source records attributed to S. Ghazaleh Dashti.

5 recordsLinked to original sources

On the use of auxiliary variables in multiple imputation when estimating the average causal effect with missing data

Estimating the average causal effect (ACE) using observational data is a key focus in causal inference for which missing data present an important challenge. Multiple imputation (MI) is a widely used method for handling missing data and can yield unbiased estimates when the imputation is compatible with the substantive analysis. One of the advantages of MI is its scope to include so-called "auxiliary variables", defined as variables associated with incomplete variables that are excluded from the substantive analysis. Although many studies have looked at the use of auxiliary variables in MI for improving precision, the study of auxiliary variables that are necessary for the identifiability (or "recoverability") of the ACE in the presence of missing data has been scant. In this work, we investigate the use of auxiliary variables, both mediators and non-mediators, across a range of typical univariable and multivariable missingness mechanisms depicted by missingness directed acyclic graphs (m-DAGs). For each setting, we derive recoverability results, then evaluate MI-based and complete-case methods for estimating the ACE using correctly specified g-computation, considering different strategies for incorporating auxiliary variables and varying degrees of compatibility for MI models. Based on findings from the simulation studies, we provide practical guidance, highlighting that distinguishing appropriately between mediator and non-mediator auxiliary variables is important to avoid bias as is the use of compatible and flexible (non-parametric) MI methods that incorporate these variables.

stat.ME↗

Handling multivariable missing data in causal mediation analysis estimating interventional effects

The interventional effects approach to causal mediation analysis is increasingly common in epidemiologic research, given its potential to address policy-relevant questions about hypothetical mediator interventions. Multiple imputation (MI) is widely used for handling missing data in epidemiologic studies. However, guidance is lacking on best practices for using MI when estimating interventional mediation effects, specifically regarding the role of the missingness mechanism in the method's performance, how to appropriately specify the MI model when g-computation is used for effect estimation, and suitable approaches to variance estimation. To address this gap, we conducted simulations based on the Victorian Adolescent Health Cohort Study. We considered seven missingness mechanisms involving varying assumptions about the influence of an intermediate confounder, a mediator, and/or the outcome on missingness in key variables. We compared the performance of complete-case analysis, six MI approaches using fully conditional specification (differing in how the imputation model was tailored), and a "substantive model compatible" multiple imputation-fully conditional specification approach. We evaluated MIBoot (MI, then bootstrap) and BootMI (bootstrap, then MI) approaches for variance estimation. All MI approaches, apart from those clearly diverging from best practice, yielded approximately unbiased estimates when none of the intermediate confounder, mediator, and outcome variables influenced missingness in any of these variables, and showed non-negligible bias otherwise. We observed the largest bias for interventional effects when each of the intermediate confounders, mediators, and outcomes influenced their own missingness. BootMI returned variance estimates with smaller bias than MIBoot.

stat.AP↗

Sensitivity analysis methods for outcome missingness using substantive-model-compatible multiple imputation and their application in causal inference

When using multiple imputation (MI) for missing data, maintaining compatibility between the imputation model and substantive analysis is important for avoiding bias. For example, some causal inference methods incorporate an outcome model with exposure-confounder interactions that must be reflected in the imputation model. Two approaches for compatible imputation with multivariable missingness have been proposed: Substantive-Model-Compatible Fully Conditional Specification (SMCFCS) and a stacked-imputation-based approach (SMC-stack). If the imputation model is correctly specified, both approaches are guaranteed to be unbiased under the "missing at random" assumption. However, this assumption is violated when the outcome causes its own missingness, which is common in practice. In such settings, sensitivity analyses are needed to assess the impact of alternative assumptions on results. An appealing solution for sensitivity analysis is delta-adjustment using MI, specifically "not-at-random" (NAR)FCS. However, the issue of imputation model compatibility has not been considered in sensitivity analysis, with a naive implementation of NARFCS being susceptible to bias. To address this gap, we propose two approaches for compatible sensitivity analysis when the outcome causes its own missingness. The proposed approaches, NAR-SMCFCS and NAR-SMC-stack, extend SMCFCS and SMC-stack, respectively, with delta-adjustment for the outcome. We evaluate these approaches using a simulation study that is motivated by a case study, to which the methods were also applied. The simulation results confirmed that a naive implementation of NARFCS produced bias in effect estimates, while NAR-SMCFCS and NAR-SMC-stack were approximately unbiased. The proposed compatible approaches provide promising avenues for conducting sensitivity analysis to missingness assumptions in causal inference.

stat.ME↗

Handling missing data when estimating causal effects with Targeted Maximum Likelihood Estimation

Targeted Maximum Likelihood Estimation (TMLE) is increasingly used for doubly robust causal inference, but how missing data should be handled when using TMLE with data-adaptive approaches is unclear. Based on the Victorian Adolescent Health Cohort Study, we conducted a simulation study to evaluate eight missing data methods in this context: complete-case analysis, extended TMLE incorporating outcome-missingness model, missing covariate missing indicator method, five multiple imputation (MI) approaches using parametric or machine-learning models. Six scenarios were considered, varying in exposure/outcome generation models (presence of confounder-confounder interactions) and missingness mechanisms (whether outcome influenced missingness in other variables and presence of interaction/non-linear terms in missingness models). Complete-case analysis and extended TMLE had small biases when outcome did not influence missingness in other variables. Parametric MI without interactions had large bias when exposure/outcome generation models included interactions. Parametric MI including interactions performed best in bias and variance reduction across all settings, except when missingness models included a non-linear term. When choosing a method to handle missing data in the context of TMLE, researchers must consider the missingness mechanism and, for MI, compatibility with the analysis method. In many settings, a parametric MI approach that incorporates interactions and non-linearities is expected to perform well.

stat.ME↗

Recoverability and estimation of causal effects under typical multivariable missingness mechanisms

In the context of missing data, the identifiability or "recoverability" of the average causal effect (ACE) depends on causal and missingness assumptions. The latter can be depicted by adding variable-specific missingness indicators to causal diagrams, creating "missingness-directed acyclic graphs" (m-DAGs). Previous research described ten canonical m-DAGs, representing typical multivariable missingness mechanisms in epidemiological studies, and determined the recoverability of the ACE in the absence of effect modification. We extend the research by determining the recoverability of the ACE in settings with effect modification and conducting a simulation study evaluating the performance of widely used missing data methods when estimating the ACE using correctly specified g-computation, which has not been previously studied. Methods assessed were complete case analysis (CCA) and various multiple imputation (MI) implementations regarding the degree of compatibility with the outcome model used in g-computation. Simulations were based on an example from the Victorian Adolescent Health Cohort Study (VAHCS), where interest was in estimating the ACE of adolescent cannabis use on mental health in young adulthood. In the canonical m-DAGs that excluded unmeasured common causes of missingness indicators, we derived the recoverable ACE if no incomplete variable causes its missingness, and non-recoverable otherwise. Besides, the simulation showed that compatible MI approaches may enable approximately unbiased ACE estimation, unless the outcome causes its missingness or it causes the missingness of a variable that causes its missingness. Researchers must consider sensitivity analysis methods incorporating external information in the latter setting. The VAHCS case study illustrates the practical implications of these findings.

stat.ME↗