SearcharxivSearch

arXiv subjects

Helene Jacqmin-Gadda

Publications and source records attributed to Helene Jacqmin-Gadda.

5 recordsLinked to original sources

Built-in Selection Bias in Proportional Hazards Models with Omitted Covariates: Simulation Evidence and Alternative Approaches

In time-to-event analysis, the hazard ratio (HR) derived from the Cox proportional hazards (PH) model is the most commonly used and widely reported measure for assessing treatment effects. However, hazard ratios are non-collapsible due to their inherent conditioning on survival up to each time point. As a result, they are subject to built-in selection bias in the presence of unmeasured heterogeneity arising from omitted important covariates, even when these covariates are independent of the main exposure at baseline, as is the case in randomized controlled trials. This article aims to provide an overview of key findings from the literature on how unobserved heterogeneity, due to omitted covariates that affect the outcome, can bias the estimation of the treatment hazard ratio in standard proportional hazards models, even in randomized trials where treatment is assigned independently of such covariates. Through simulations, we evaluate the extent of bias in the semi-parametric Cox PH model and parametric PH model under various scenarios of unmeasured heterogeneity. We then compare these standard models to alternative approaches that either account for this issue or are considered robust to it. These alternatives include the hazard ratio estimated from frailty models, regression parameters from an Accelerated Failure Time (AFT) model, and survival differences between treatment groups estimated nonparametrically using Kaplan-Meier curves or based on a Cox model with time-dependent effect of the exposure. We illustrate the practical relevance of the explored alternatives through a real data application to a randomized controlled trial from the Radiation Therapy Oncology Group (RTOG 9202).

stat.ME

A Two-Stage Bayesian Approach for Variable Selection in Joint Modeling of Multiple Longitudinal Markers with Competing Risks

In many clinical and epidemiological studies, collecting longitudinal measurements together with time-to-event outcomes is essential. Accurately estimating the association between longitudinal markers and event risks, as well as identifying key markers for prediction, is especially important in the presence of competing risks. However, as the number of markers increases, fitting full joint models becomes computationally difficult and may lead to convergence issues. We propose a two-stage Bayesian approach for variable selection in joint models with multiple longitudinal markers and competing risks. The method efficiently identifies important longitudinal markers and covariates. In the first stage, a one-marker joint model is fitted for each marker with the competing risks outcome, and individual marker trajectories are predicted, reducing bias from informative dropout. In the second stage, a cause-specific hazards model is fitted, incorporating the predicted current values of all markers as time-dependent covariates. We consider both continuous and Dirac spike-and-slab priors for Bayesian variable selection, implemented through MCMC algorithms. Our approach enables risk prediction using a large number of longitudinal markers, which is often infeasible for standard joint models. We evaluate performance through simulation studies, examining both variable selection and predictive accuracy. Finally, we apply the method to predict dementia risk in the Three-City (3C) study, a French cohort with competing risks of death. To facilitate use, we provide an R package, VSJM, available at: https:/github.com/tbaghfalaki/VSJM.

stat.ME

Dynamic prediction of an event using multiple longitudinal markers: a model averaging approach

Dynamic event prediction, using joint modeling of survival time and longitudinal variables, is extremely useful in personalized medicine. However, the estimation of joint models including many longitudinal markers is still a computational challenge because of the high number of random effects and parameters to be estimated. In this paper, we propose a model averaging strategy to combine predictions from several joint models for the event, including one longitudinal marker only or pairwise longitudinal markers. The prediction is computed as the weighted mean of the predictions from the one-marker or two-marker models, with the time-dependent weights estimated by minimizing the time-dependent Brier score. This method enables us to combine a large number of predictions issued from joint models to achieve a reliable and accurate individual prediction. Advantages and limits of the proposed methods are highlighted in a simulation study by comparison with the predictions from well-specified and misspecified all-marker joint models as well as the one-marker and two-marker joint models. Using the PBC2 data set, the method is used to predict the risk of death in patients with primary biliary cirrhosis. The method is also used to analyze a French cohort study called the 3C data. In our study, seventeen longitudinal markers are considered to predict the risk of death.

stat.ME

A Two-stage Joint Modeling Approach for Multiple Longitudinal Markers and Time-to-event Data

Collecting multiple longitudinal measurements and time-to-event outcomes is a common practice in clinical and epidemiological studies, often focusing on exploring associations between them. Joint modeling is the standard analytical tool for such data, with several R packages available. However, as the number of longitudinal markers increases, the computational burden and convergence challenges make joint modeling increasingly impractical. This paper introduces a novel two-stage Bayesian approach to estimate joint models for multiple longitudinal measurements and time-to-event outcomes. The method builds on the standard two-stage framework but improves the initial stage by estimating a separate one-marker joint model for the event and each longitudinal marker, rather than relying on mixed models. These estimates are used to derive predictions of individual marker trajectories, avoiding biases from informative dropouts. In the second stage, a proportional hazards model is fitted, incorporating the predicted current values and slopes of the markers as time-dependent covariates. To address uncertainty in the first-stage predictions, a multiple imputation technique is employed when estimating the Cox model in the second stage. This two-stage method allows for the analysis of numerous longitudinal markers, which is often infeasible with traditional multi-marker joint modeling. The paper evaluates the approach through simulation studies and applies it to the PBC2 dataset and a real-world dementia dataset containing 17 longitudinal markers. An R package, TSJM, implementing the method is freely available on GitHub: https://github.com/tbaghfalaki/TSJM.

stat.ME

A Newton-Like Algorithm for Likelihood Maximization: The Robust-Variance Scoring Algorithm

This article studies a Newton-like method already used by several authors but which has not been thouroughly studied yet. We call it the robust-variance scoring (RVS) algorithm because the main version of the algorithm that we consider replaces minus the Hessian of the loglikelihood used in the Newton-Raphson algorithm by a matrix $G$ which is an estimate of the variance of the score under the true probability, which uses only the individual scores. Thus an iteration of this algorithm requires much less computations than an iteration of the Newton-Raphson algorithm. Moreover this estimate of the variance of the score estimates the information matrix at maximum. We have also studied a convergence criterion which has the nice interpretation of estimating the ratio of the approximation error over the statistical error; thus it can be used for stopping the iterative process whatever the problem. A simulation study confirms that the RVS algorithm is faster than the Marquardt algorithm (a robust version of the Newton-Raphson algorithm); this happens because the number of iterations needed by the RVS algorithm is barely larger than that needed by the Marquardt algorithm while the computation time for each iteration is much shorter. Also the coverage rates using the matrix $G$ are satisfactory.

math.ST