SearcharxivSearch

arXiv subjects

Denis Talbot

Publications and source records attributed to Denis Talbot.

13 recordsLinked to original sources

Causal inference via propensity scores for case-control studies

Propensity score methods for causal inference are increasingly being used in cohort and experimental designs, but their development and uptake in outcome-dependent sampling schemes, such as case-control studies, remains limited. Case-control studies involve the sampling of individuals with and without an outcome of interest with the goal of estimating the effects of past exposures. When the design is observational, statistical adjustment for confounding bias is necessary. In case-control studies, propensity score models can be fit using control data under the assumption that the controls are representative of the source population with respect to their exposure distribution conditional on covariates ("control exchangeability"). In this paper, we first demonstrate that relative effects, such as causal risk ratios, are estimable under three different case-control design variants using control-fitted propensity scores. We appropriate two existing estimators for these designs: inverse probability of treatment weighting and an efficient and doubly robust estimator. We also introduce a novel two-step propensity score caliper-matching procedure for case-control designs. We introduce novel diagnostic tools to verify two necessary types of overlap. We then contrast our estimators using simulated data and apply them to examine the association between regular aspirin use and ovarian cancer risk.

stat.ME

The positivity assumption in causal mediation analyses? Checked!

Causal mediation analyses are increasingly used in psychological sciences. Among the required assumptions, positivity is unfortunately seldom mentioned, likely due to the lack of tools for checking it. Mediational positivity is more complex than positivity in standard, non-mediated exposure effect analysis, because it requires positivity for both the exposure and the mediator, and because the specific form of the positivity assumption depends on the mediation estimand of interest. We propose an extension of the Positivity Regression Trees (PoRT) algorithm -- which was recently designed to check positivity in non-mediated settings without requiring assumptions about the modeling or the data-generating process -- to controlled, natural and interventional mediational effects. We illustrate its use through an application in mental health and have made it accessible through the port R package and a related notebook available at github.com/ArthurChatton/dePoRT-notebook. Finally, we discuss consequences and provide recommendations for when positivity violations are identified.

stat.ME

trajmsm: An R package for Trajectory Analysis and Causal Modeling

The R package trajmsm provides functions designed to simplify the estimation of the parameters of a model combining latent class growth analysis (LCGA), a trajectory analysis technique, and marginal structural models (MSMs) called LCGA-MSM. LCGA summarizes similar patterns of change over time into a few distinct categories called trajectory groups, which are then included as "treatments" in the MSM. MSMs are a class of causal models that correctly handle treatment-confounder feedback. The parameters of LCGA-MSMs can be consistently estimated using different estimators, such as inverse probability weighting (IPW), g-computation, and pooled longitudinal targeted maximum likelihood estimation (pooled LTMLE). These three estimators of the parameters of LCGA-MSMs are currently implemented in our package. In the context of a time-dependent outcome, we previously proposed a combination of LCGA and history-restricted MSMs (LCGA-HRMSMs). Our package provides additional functions to estimate the parameters of such models. Version 0.1.3 of the package is currently available on CRAN.

stat.ME

Adaptive sparsening and smoothing of the treatment model for longitudinal causal inference using outcome-adaptive LASSO and marginal fused LASSO

Causal variable selection in time-varying treatment settings is challenging due to evolving confounding effects. Existing methods mainly focus on time-fixed exposures and are not directly applicable to time-varying scenarios. We propose a novel two-step procedure for variable selection when modeling the treatment probability at each time point. We first introduce a novel approach to longitudinal confounder selection using a Longitudinal Outcome Adaptive LASSO (LOAL) that will data-adaptively select covariates with theoretical justification of variance reduction of the estimator of the causal effect. We then propose an Adaptive Fused LASSO that can collapse treatment model parameters over time points with the goal of simplifying the models in order to improve the efficiency of the estimator while minimizing model misspecification bias compared with naive pooled logistic regression models. Our simulation studies highlight the need for and usefulness of the proposed approach in practice. We implemented our method on data from the Nicotine Dependence in Teens study to estimate the effect of the timing of alcohol initiation during adolescence on depressive symptoms in early adulthood.

stat.ME

Efficient adjustment sets for time-dependent treatment effect estimation in nonparametric causal graphical model

Criteria for identifying optimal adjustment sets yielding consistent estimation with minimal asymptotic variance of average treatment effects in parametric and nonparametric models have recently been established. In a single treatment time point setting, it has been shown that the optimal adjustment set can be identified based on a causal directed acyclic graph alone. In a time-dependent treatment setting, previous work has established graphical rules to compare the asymptotic variance of estimators based on nested time-dependent adjustment sets. However, these rules do not always permit the identification of an optimal time-dependent adjustment set based on a causal graph alone. We extend those results by exploiting conditional independencies that can be read from the graph and demonstrate theoretically and empirically that our results can yield estimators with lower asymptotic variance than those allowed by previous results. We further show how our results allow for the identification of optimal adjustment sets based on a directed acyclic graph alone in the time-dependent treatment setting.

math.ST

Evaluation and comparison of covariate balance metrics in studies with time-dependent confounding

Marginal structural models have been increasingly used by analysts in recent years to account for confounding bias in studies with time-varying treatments. The parameters of these models are often estimated using inverse probability of treatment weighting. To ensure that the estimated weights adequately control confounding, it is possible to check for residual imbalance between treatment groups in the weighted data. Several balance metrics have been developed and compared in the cross-sectional case but have not yet been evaluated and compared in longitudinal studies with time-varying treatment. We have first extended the definition of several balance metrics to the case of a time-varying treatment, with or without censoring. We then compared the performance of these balance metrics in a simulation study by assessing the strength of the association between their estimated level of imbalance and bias. We found that the Mahalanobis balance performed best.Finally, the method was illustrated for estimating the cumulative effect of statins exposure over one year on the risk of cardiovascular disease or death in people aged 65 and over in population-wide administrative data. This illustration confirms the feasibility of employing our proposed metrics in large databases with multiple time-points.

stat.ME

A Double Machine Learning Approach for the Evaluation of COVID-19 Vaccine Effectiveness under the Test-Negative Design: Analysis of Qu\'ebec Administrative Data

The test-negative design (TND), which is routinely used for monitoring seasonal flu vaccine effectiveness (VE), has recently become integral to COVID-19 vaccine surveillance, notably in Qu\'ebec, Canada. Some studies have addressed the identifiability and estimation of causal parameters under the TND, but efficiency bounds for nonparametric estimators of the target parameter under the unconfoundedness assumption have not yet been investigated. Motivated by the goal of improving adjustment for measured confounders when estimating COVID-19 VE among community-dwelling people aged $\geq 60$ years in Qu\'ebec, we propose a one-step doubly robust and locally efficient estimator called TNDDR (TND doubly robust), which utilizes cross-fitting (sample splitting) and can incorporate machine learning techniques to estimate the nuisance functions and thus improve control for measured confounders. We derive the efficient influence function (EIF) for the marginal expectation of the outcome under a vaccination intervention, explore the von Mises expansion, and establish the conditions for $\sqrt{n}-$consistency, asymptotic normality and double robustness of TNDDR. The proposed estimator is supported by both theoretical and empirical justifications.

stat.ME

Estimation of the attributable fraction for time to event outcomes using an inverse probability of exposure weighted Kaplan-Meier estimator

Population attributable fractions aim to quantify the proportion of the cases of an outcome (for example, a disease) that would have been avoided had no individuals in the population been exposed to a given exposure. This quantity thus plays a crucial role in epidemiology and public health, notably to guide policies, interventions or to assess the burden of a disease due to a particular exposure. Various statistical methods have been proposed to estimate attributable fractions using observational data. When time-to-event data are used, several of these formulas yield invalid results. Alternative valid formulas are available but remain scarcely used. We propose a new estimator of the attributable fraction that is both conceptually simple and easy to implement using common statistical software. Our proposed estimator makes use of the Kaplan-Meier estimator to address censoring and potentially non-proportional hazards, as well as inverse probability weighting to control confounding. Nonparametric bootstrap is proposed to produce inferences. A simulation study is used to illustrate and compare our proposed estimator to several alternatives. The results showcase the bias of many commonly used traditional approaches and the validity of our estimator under its working assumptions.

stat.ME

A bootstrap approach for validating the number of groups identified by latent class growth models

The use of longitudinal finite mixture models such as group-based trajectory modeling has seen a sharp increase during the last decades in the medical literature. However, these methods have been criticized especially because of the data-driven modelling process which involves statistical decision-making. In this paper, we propose an approach that uses bootstrap to sample observations with replacement from the original data to validate the number of groups identified and to quantify the uncertainty in the number of groups. The method allows investigating the statistical validity and the uncertainty of the groups identified in the original data by checking if the same solution is also found across the bootstrap samples. In a simulation study, we examined whether the bootstrap-estimated variability in the number of groups reflected the replication-wise variability. We also compared the replication-wise variability to the Bayesian posterior probability. We evaluated the ability of three commonly used adequacy criteria (average posterior probability, odds of correct classification and relative entropy) to identify uncertainty in the number of groups. Finally, we illustrated the proposed approach using data from the Quebec Integrated Chronic Disease Surveillance System to identify longitudinal medication patterns between 2015 and 2018 in older adults with diabetes.

stat.ME

An Alternative Perspective on the Robust Poisson Method for Estimating Risk or Prevalence Ratios

The robust Poisson method is becoming increasingly popular when estimating the association of exposures with a binary outcome. Unlike the logistic regression model, the robust Poisson method yields results that can be interpreted as risk or prevalence ratios. In addition, it does not suffer from frequent non-convergence problems like the most common implementations of maximum likelihood estimators of the log-binomial model. However, using a Poisson distribution to model a binary outcome may seem counterintuitive. Methodological papers have often presented this as a good approximation to the more natural binomial distribution. In this paper, we provide an alternative perspective to the robust Poisson method based on the semiparametric theory. This perspective highlights that the robust Poisson method does not require assuming a Poisson distribution for the outcome. In fact, the method only assumes a log-linear relationship between the risk/prevalence of the outcome and the explanatory variables. This assumption and consequences of its violation are discussed. Suggestions to reduce the risk of violating the modeling assumption are also provided. Additionally, we discuss and contrast the robust Poisson method with other approaches for estimating exposure risk or prevalence ratios.

stat.ME

Double robust estimation of partially adaptive treatment strategies

Precision medicine aims to tailor treatment decisions according to patients' characteristics. G-estimation and dynamic weighted ordinary least squares (dWOLS) are double robust statistical methods that can be used to identify optimal adaptive treatment strategies. They require both a model for the outcome and a model for the treatment and are consistent if at least one of these models is correctly specified. It is underappreciated that these methods additionally require modeling all existing treatment-confounder interactions to yield consistent estimators. Identifying partially adaptive treatment strategies that tailor treatments according to only a few covariates, ignoring some interactions, may be preferable in practice. It has been proposed to combine inverse probability weighting and G-estimation to address this issue, but we argue that the resulting estimator is not expected to be double robust. Building on G-estimation and dWOLS, we propose alternative estimators of partially adaptive strategies and demonstrate their double robustness. We investigate and compare the empirical performance of six estimators in a simulation study. As expected, estimators combining inverse probability weighting with either G-estimation or dWOLS are biased when the treatment model is incorrectly specified. The other estimators are unbiased if either the treatment or the outcome model are correctly specified and have similar standard errors. Using data maintained by the Centre des Maladies du Sein, the methods are illustrated to estimate a partially adaptive treatment strategy for tailoring hormonal therapy use in breast cancer patients according to their estrogen receptor status and body mass index. R software implementing our estimators is provided.

stat.ME

Marginal structural models with Latent Class Growth Modeling of Treatment Trajectories

In a real-life setting, little is known regarding the effectiveness of statins for primary prevention among older adults, and analysis of observational data can add crucial information on the benefits of actual patterns of use. Latent class growth models (LCGM) are increasingly proposed as a solution to summarize the observed longitudinal treatment in a few distinct groups. When combined with standard approaches like Cox proportional hazards models, LCGM can fail to control time-dependent confounding bias because of time-varying covariates that have a double role of confounders and mediators. We propose to use LCGM to classify individuals into a few latent classes based on their medication adherence pattern, then choose a working marginal structural model (MSM) that relates the outcome to these groups. The parameter of interest is nonparametrically defined as the projection of the true MSM onto the chosen working model. The combination of LCGM with MSM is a convenient way to describe treatment adherence and can effectively control time-dependent confounding. Simulation studies were used to illustrate our approach and compare it with unadjusted, baseline covariates-adjusted, time-varying covariates adjusted and inverse probability of trajectory groups weighting adjusted models. We found that our proposed approach yielded estimators with little or no bias.

stat.ME

A generalized double robust Bayesian model averaging approach to causal effect estimation with application to the Study of Osteoporotic Fractures

Analysts often use data-driven approaches to supplement their substantive knowledge when selecting covariates for causal effect estimation. Multiple variable selection procedures tailored for causal effect estimation have been devised in recent years, but additional developments are still required to adequately address the needs of data analysts. In this paper, we propose a Generalized Bayesian Causal Effect Estimation (GBCEE) algorithm to perform variable selection and produce double robust estimates of causal effects for binary or continuous exposures and outcomes. GBCEE employs a prior distribution that targets the selection of true confounders and predictors of the outcome for the unbiased estimation of causal effects with reduced standard errors. Double robust estimators provide some robustness against model misspecification, whereas the Bayesian machinery allows GBCEE to directly produce inferences for its estimate. GBCEE was compared to multiple alternatives in various simulation scenarios and was observed to perform similarly or to outperform double robust alternatives. Its ability to directly produce inferences is also an important advantage from a computational perspective. The method is finally illustrated for the estimation of the effect of meeting physical activity recommendations on the risk of hip or upper-leg fractures among elderly women in the Study of Osteoporotic Fractures. The 95% confidence interval produced by GBCEE is 61% shorter than that of a double robust estimator adjusting for all potential confounders in this illustration.

stat.ME