Searcharxiv⌕ Search

arXiv subjects

Jonathan W. Bartlett

Publications and source records attributed to Jonathan W. Bartlett.

At least 19 recordsLinked to original sources

Bayesian analysis of the causal reference-based model for missing data in clinical trials, accommodating partially observed post-intercurrent event data

When treatment policy estimands are of interest, clinical trials often attempt to collect patient data after intercurrent events (ICEs), although such data are often limited. Retrieved dropout imputation methods, which use pre-ICE and available post-ICE data to impute missing post-ICE outcomes, are commonly applied but often yield treatment effect estimates with large standard errors (SEs) and may encounter convergence issues when post-ICE data are sparse. Reference-based imputation methods are also used, but they rely on strong assumptions about post-ICE outcomes, which can lead to biased estimates if these assumptions are incorrect. To address these limitations, we previously proposed the reference-based Bayesian causal model (BCM), which incorporates a prior on the maintained effect parameter to reflect uncertainty in reference-based assumptions for missing post-ICE data. Our earlier work assumed no post-ICE data were observed. Here, we extend the BCM to incorporate available post-ICE outcomes, providing an approach that mitigates limitations of both retrieved-dropout and standard reference-based methods. We propose both a fully Bayesian model and an imputation-based approach. A simulation study was conducted to evaluate the frequentist properties of the proposed methods in settings with partially observed post-ICE data and to compare performance with existing approaches. Retrieved-dropout methods produced higher estimated SEs than the BCM, particularly when post-ICE data were sparse. Under the BCM, treatment effect SEs increased as post-ICE data became more limited for both modelling approaches. Importantly, this increase can be controlled through the prior variance of the maintained effect parameter, with more informative priors stabilising estimation when post-ICE data are scarce.

stat.ME↗

How to interpret hazard ratios

The hazard ratio, typically estimated using Cox's famous proportional hazards model, is the most common effect measure used to describe the association or effect of a covariate on a time-to-event outcome. In recent years the hazard ratio has been argued by some to lack a causal interpretation, even in randomised trials, and even if the proportional hazards assumption holds. This is concerning, not least due to the ubiquity of hazard ratios in analyses of time-to-event data. We review these criticisms, describe how we think hazard ratios should be interpreted, and argue that they retain a valid causal interpretation. Nevertheless, alternative measures may be preferable to describe effects of exposures or treatments on time-to-event outcomes.

stat.ME↗

Comparison of Parametric versus Machine-learning Multiple Imputation in Clinical Trials with Missing Continuous Outcomes

The use of flexible machine-learning (ML) models to generate imputations of missing data within the framework of Multiple Imputation (MI) has recently gained traction, particularly in observational settings. For randomised controlled trials (RCTs), it is unclear whether ML approaches to MI provide valid inference, and whether they outperform parametric MI approaches under complex data generating mechanisms. We conducted two simulations in RCT settings that have incomplete continuous outcomes but fully observed covariates. We compared Complete Cases, standard MI (MI-norm), MI with predictive mean matching (MI-PMM) and ML-based approaches to MI, including classification and regression trees (MI-CART), Random Forests (MI-RF) and SuperLearner when outcomes are missing completely at random or missing at random conditional on treatment/covariate. The first simulation explored non-linear covariate-outcome relationships in the presence/absence of covariate-treatment interactions. The second simulation explored skewed repeated measures, motivated by a trial with digital outcomes. In the absence of interactions, we found that Complete Cases yields reliable inference; MI-norm performs similarly, except when missingness depends on the covariate. ML approaches can lead to smaller mean squared error than Complete Cases and MI-norm in specific non-linear settings, but provide unreliable inference for others. MI-PMM can lead to unreliable inference in several settings. In the presence of complex treatment-covariate interactions, performing MI separately by arm, either with MI-norm, MI-RF or MI-CART, provides inference that has comparable or with better properties compared to Complete Cases when the analysis model omits the interaction. For ML approaches, we observed unreliable inference in terms of bias in the estimated effect and/or its standard error when Rubin's Rules are implemented.

stat.ME↗

The role of post intercurrent event data in the estimation of hypothetical estimands in clinical trials

Estimation of hypothetical estimands in clinical trials typically does not make use of data that may be collected after the intercurrent event (ICE). Some recent papers have shown that such data can be used for estimation of hypothetical estimands, and that statistical efficiency and power can be increased compared to using estimators that only use data before the ICE. In this paper we critically examine the efficiency and bias of estimators that do and do not exploit data collected after ICEs, in a simplified setting. We find that efficiency can only be improved by assuming certain covariate effects are common between patients who do and do not experience ICEs, and that even when such an assumption holds, gains in efficiency will typically be modest. We moreover argue that the assumptions needed to gain efficiency by using post-ICE outcomes will often not hold, such that estimators using post-ICE data may lead to biased estimates and invalid inferences. As such, we recommend that in general estimation of hypothetical estimands should be based on estimators that do not make use of post-ICE data.

stat.ME↗

Dealing with multiple intercurrent events using hypothetical and treatment policy strategies simultaneously

To precisely define the treatment effect of interest in a clinical trial, the ICH E9 estimand addendum describes that relevant so-called intercurrent events should be identified and strategies specified to deal with them. Handling intercurrent events with different strategies leads to different estimands. In this paper, we focus on estimands that involve addressing one intercurrent event with the treatment policy strategy and another with the hypothetical strategy. We define these estimands using potential outcomes and causal diagrams, considering the possible causal relationships between the two intercurrent events and other variables. We show that there are different causal estimand definitions and assumptions one could adopt, each having different implications for estimation, which is demonstrated in a simulation study. The different considerations are illustrated conceptually using a diabetes trial as an example.

stat.ME↗

The Estimand Framework and Causal Inference: Complementary not Competing Paradigms

The creation of the ICH E9 (R1) estimands framework has led to more precise specification of the treatment effects of interest in the design and statistical analysis of clinical trials. However, it is unclear how the new framework relates to causal inference, as both approaches appear to define what is being estimated and have a quantity labelled an estimand. Using illustrative examples, we show that both approaches can be used to define a population-based summary of an effect on an outcome for a specified population and highlight the similarities and differences between these approaches. We demonstrate that the ICH E9 (R1) estimand framework offers a descriptive, structured approach that is more accessible to non-mathematicians, facilitating clearer communication of trial objectives and results. We then contrast this with the causal inference framework, which provides a mathematically precise definition of an estimand, and allows the explicit articulation of assumptions through tools such as causal graphs. Despite these differences, the two paradigms should be viewed as complementary rather than competing. The combined use of both approaches enhances the ability to communicate what is being estimated. We encourage those familiar with one framework to appreciate the concepts of the other to strengthen the robustness and clarity of clinical trial design, analysis, and interpretation.

stat.ME↗

Sensitivity analysis methods for outcome missingness using substantive-model-compatible multiple imputation and their application in causal inference

When using multiple imputation (MI) for missing data, maintaining compatibility between the imputation model and substantive analysis is important for avoiding bias. For example, some causal inference methods incorporate an outcome model with exposure-confounder interactions that must be reflected in the imputation model. Two approaches for compatible imputation with multivariable missingness have been proposed: Substantive-Model-Compatible Fully Conditional Specification (SMCFCS) and a stacked-imputation-based approach (SMC-stack). If the imputation model is correctly specified, both approaches are guaranteed to be unbiased under the "missing at random" assumption. However, this assumption is violated when the outcome causes its own missingness, which is common in practice. In such settings, sensitivity analyses are needed to assess the impact of alternative assumptions on results. An appealing solution for sensitivity analysis is delta-adjustment using MI, specifically "not-at-random" (NAR)FCS. However, the issue of imputation model compatibility has not been considered in sensitivity analysis, with a naive implementation of NARFCS being susceptible to bias. To address this gap, we propose two approaches for compatible sensitivity analysis when the outcome causes its own missingness. The proposed approaches, NAR-SMCFCS and NAR-SMC-stack, extend SMCFCS and SMC-stack, respectively, with delta-adjustment for the outcome. We evaluate these approaches using a simulation study that is motivated by a case study, to which the methods were also applied. The simulation results confirmed that a naive implementation of NARFCS produced bias in effect estimates, while NAR-SMCFCS and NAR-SMC-stack were approximately unbiased. The proposed compatible approaches provide promising avenues for conducting sensitivity analysis to missingness assumptions in causal inference.

stat.ME↗

Multiple imputation of missing covariates when using the Fine-Gray model

The Fine-Gray model for the subdistribution hazard is commonly used for estimating associations between covariates and competing risks outcomes. When there are missing values in the covariates included in a given model, researchers may wish to multiply impute them. Assuming interest lies in estimating the risk of only one of the competing events, this paper develops a substantive-model-compatible multiple imputation approach that exploits the parallels between the Fine-Gray model and the standard (single-event) Cox model. In the presence of right-censoring, this involves first imputing the potential censoring times for those failing from competing events, and thereafter imputing the missing covariates by leveraging methodology previously developed for the Cox model in the setting without competing risks. In a simulation study, we compared the proposed approach to alternative methods, such as imputing compatibly with cause-specific Cox models. The proposed method performed well (in terms of estimation of both subdistribution log hazard ratios and cumulative incidences) when data were generated assuming proportional subdistribution hazards, and performed satisfactorily when this assumption was not satisfied. The gain in efficiency compared to a complete-case analysis was demonstrated in both the simulation study and in an applied data example on competing outcomes following an allogeneic stem cell transplantation. For individual-specific cumulative incidence estimation, assuming proportionality on the correct scale at the analysis phase appears to be more important than correctly specifying the imputation procedure used to impute the missing covariates.

stat.ME↗

G-formula for causal inference via multiple imputation

G-formula is a popular approach for estimating treatment or exposure effects from longitudinal data that are subject to time-varying confounding. G-formula estimation is typically performed by Monte-Carlo simulation, with non-parametric bootstrapping used for inference. We show that G-formula can be implemented by exploiting existing methods for multiple imputation (MI) for synthetic data. This involves using an existing modified version of Rubin's variance estimator. In practice missing data is ubiquitous in longitudinal datasets. We show that such missing data can be readily accommodated as part of the MI procedure when using G-formula, and describe how MI software can be used to implement the approach. We explore its performance using a simulation study and an application from cystic fibrosis.

stat.ME↗

Estimating hypothetical estimands with causal inference and missing data estimators in a diabetes trial

The recently published ICH E9 addendum on estimands in clinical trials provides a framework for precisely defining the treatment effect that is to be estimated, but says little about estimation methods. Here we report analyses of a clinical trial in type 2 diabetes, targeting the effects of randomised treatment, handling rescue treatment and discontinuation of randomised treatment using the so-called hypothetical strategy. We show how this can be estimated using mixed models for repeated measures, multiple imputation, inverse probability of treatment weighting, G-formula and G-estimation. We describe their assumptions and practical details of their implementation using packages in R. We report the results of these analyses, broadly finding similar estimates and standard errors across the estimators. We discuss various considerations relevant when choosing an estimation approach, including computational time, how to handle missing data, whether to include post intercurrent event data in the analysis, whether and how to adjust for additional time-varying confounders, and whether and how to model different types of ICE separately.

stat.AP↗

Weighted hazard ratio estimation for delayed and diminishing treatment effect

Non-proportional hazards (NPH) have been observed in confirmatory clinical trials with time to event outcomes. Under NPH, the hazard ratio does not stay constant over time and the log-rank test is no longer the most powerful test. The weighted log-rank test (WLRT) has been introduced to deal with the presence of non-proportionality. We focus our attention on the WLRT and the complementary Cox model based on the time-varying treatment effect proposed by Lin and León (2017) (doi: 10.1016/j.conctc.2017.09.004). We will investigate whether the proposed weighted hazard ratio (WHR) approach is unbiased in scenarios where the WLRT statistic is the most powerful test. In the diminishing treatment effect scenario where the WLRT statistic would be most optimal, the time-varying treatment effect estimated by the Cox model estimates the treatment effect very close to the true one. However, when the true hazard ratio is large we note that the proposed model overestimates the treatment effect and the treatment profile over time. However, in the delayed treatment scenario, the estimated treatment effect profile over time is typically close to the true profile. For both scenarios, we have demonstrated analytically that the hazard ratio functions are approximately equal under certain constraints. In conclusion, our results demonstrate that in certain scenarios where a given WLRT would be most powerful, we observe that the WHR from the corresponding Cox model is estimating the treatment effect close to the true one.

stat.ME↗

Standard and reference-based conditional mean imputation

Clinical trials with longitudinal outcomes typically include missing data due to missed assessments or structural missingness of outcomes after intercurrent events handled with a hypothetical strategy. Approaches based on Bayesian random multiple imputation and Rubin's rule for pooling results across multiple imputed datasets are increasingly used in order to align the analysis of these trials with the targeted estimand. We propose and justify deterministic conditional mean imputation combined with the jackknife for inference as an alternative approach. The method is applicable to imputations under a missing-at-random assumption as well as for reference-based imputation approaches. In an application and a simulation study, we demonstrate that it provides consistent treatment effect estimates with the Bayesian approach and reliable frequentist inference with accurate standard error estimation and type I error control. A further advantage of the method is that it does not rely on random sampling and is therefore replicable and unaffected by Monte Carlo error.

stat.ME↗

Hypothetical estimands in clinical trials: a unification of causal inference and missing data methods

The ICH E9 addendum introduces the term intercurrent event to refer to events that happen after randomisation and that can either preclude observation of the outcome of interest or affect its interpretation. It proposes five strategies for handling intercurrent events to form an estimand but does not suggest statistical methods for estimation. In this paper we focus on the hypothetical strategy, where the treatment effect is defined under the hypothetical scenario in which the intercurrent event is prevented. For its estimation, we consider causal inference and missing data methods. We establish that certain 'causal inference estimators' are identical to certain 'missing data estimators'. These links may help those familiar with one set of methods but not the other. Moreover, using potential outcome notation allows us to state more clearly the assumptions on which missing data methods rely to estimate hypothetical estimands. This helps to indicate whether estimating a hypothetical estimand is reasonable, and what data should be used in the analysis. We show that hypothetical estimands can be estimated by exploiting data after intercurrent event occurrence, which is typically not used. We also present Monte Carlo simulations that illustrate the implementation and performance of the methods in different settings.

stat.ME↗

Reference based multiple imputation -- what is the right variance and how to estimate it

Reference based multiple imputation methods have become popular for handling missing data in randomised clinical trials. Rubin's variance estimator is well known to be biased compared to the reference based imputation estimator's true repeated sampling variance. Somewhat surprisingly given the increasingly popularity of these methods, there has been relatively little debate in the literature as to whether Rubin's variance estimator or alternative (smaller) variance estimators targeting the repeated sampling variance are more appropriate. We review the arguments made on both sides of this debate, and conclude that the repeated sampling variance is more appropriate. We review different approaches for estimating the frequentist variance, and suggest a recent proposal for combining bootstrapping with multiple imputation as a widely applicable general solution. At the same time, in light of the consequences of reference based assumptions for frequentist variance, we believe further scrutiny of these methods is warranted to determine whether the the strength of their assumptions are generally justifiable.

stat.ME↗

Bootstrap Inference for Multiple Imputation under Uncongeniality and Misspecification

Multiple imputation has become one of the most popular approaches for handling missing data in statistical analyses. Part of this success is due to Rubin's simple combination rules. These give frequentist valid inferences when the imputation and analysis procedures are so called congenial and the complete data analysis is valid, but otherwise may not. Roughly speaking, congeniality corresponds to whether the imputation and analysis models make different assumptions about the data. In practice imputation and analysis procedures are often not congenial, such that tests may not have the correct size and confidence interval coverage deviates from the advertised level. We examine a number of recent proposals which combine bootstrapping with multiple imputation, and determine which are valid under uncongeniality and model misspecification. Imputation followed by bootstrapping generally does not result in valid variance estimates under uncongeniality or misspecification, whereas bootstrapping followed by imputation does. We recommend a particular computationally efficient variant of bootstrapping followed by imputation.

stat.ME↗

Measurement error as a missing data problem

This article focuses on measurement error in covariates in regression analyses in which the aim is to estimate the association between one or more covariates and an outcome, adjusting for confounding. Error in covariate measurements, if ignored, results in biased estimates of parameters representing the associations of interest. Studies with variables measured with error can be considered as studies in which the true variable is missing, for either some or all study participants. We make the link between measurement error and missing data and describe methods for correcting for bias due to covariate measurement error with reference to this link, including regression calibration (conditional mean imputation), maximum likelihood and Bayesian methods, and multiple imputation. The methods are illustrated using data from the Third National Health and Nutrition Examination Survey (NHANES III) to investigate the association between the error-prone covariate systolic blood pressure and the hazard of death due to cardiovascular disease, adjusted for several other variables including those subject to missing data. We use multiple imputation and Bayesian approaches that can address both measurement error and missing data simultaneously. Example data and R code are provided in supplementary materials.

stat.ME↗

Robustness of ANCOVA in randomised trials with unequal randomisation

Randomised trials with continuous outcomes are often analysed using ANCOVA, with adjustment for prognostic baseline covariates. In an article published recently, Wang \etal proved that in this setting the model based standard error estimator for the treamtent effect is consistent under outcome model misspecification, provided the probability of randomisation to each treatment is 1/2. In this article, we extend their results allowing for unequal randomisation. These demonstrate that the model based standard error is in general inconsistent when the randomisation probability differs from 1/2. In contrast, the sandwich standard error can provide asymptotically valid inferences under misspecification when randomisation probabilities are not equal, and is therefore recommended when randomisation is unequal.

stat.ME↗

Covariate adjustment and prediction of mean response in randomised trials

Analyses of randomised trials are often based on regression models which adjust for baseline covariates, in addition to randomised group. Based on such models, one can obtain estimates of the marginal mean outcome for the population under assignment to each treatment, by averaging the model based predictions across the empirical distribution of the baseline covariates in the trial. We identify under what conditions such estimates are consistent, and in particular show that for canonical generalised linear models, the resulting estimates are always consistent. We show that a recently proposed variance estimator underestimates the true variance when the baseline covariates are not fixed in repeated sampling, and provide a simple adjustment to remedy this. We also describe an alternative semiparametric estimator which is consistent even when the outcome regression model used is misspecified. The different estimators are compared through simulations and application to a recently conducted trial in asthma.

stat.ME↗