Searcharxiv⌕ Search

arXiv subjects

Nan van Geloven

Publications and source records attributed to Nan van Geloven.

9 recordsLinked to original sources

Population-Level Decision Curve Analysis May Mislead the Evaluation of Prediction Model Usefulness under Subgroup Utility Heterogeneity

Background Decision curve analysis (DCA) evaluates prediction model usefulness using net benefit (NB). Population-level NB is often interpreted as a proxy for population-level expected utility, but this assumes comparability of utility across subgroups. We aimed to characterize when subgroup utility heterogeneity may invalidate population-level DCA conclusions. Methods We compared prediction-driven treatment with default strategies in populations containing subgroups with different utility values. We derived an inconsistency region: combinations of subgroup-specific ΔNB values for which population-level NB and utility favor different strategies. We also developed a practical robustness framework. Results Opposite signs of subgroup-specific ΔNB provide a warning signal for possible inconsistency. The inconsistency region is larger when subgroup sizes are more similar and subgroup-specific \(a-c\) values are more different, where \(a-c\) is the incremental utility of a true positive relative to a false negative. When \(a-c\) differs across subgroups, population-level NB combines quantities on different implicit utility scales and may conflict with population utility. A real-world case study illustrates the problem. Conclusions Using population-level NB as a proxy for population utility implicitly assumes homogeneous \(a-c\) across subgroups. Subgroup DCA and our framework can identify and assess when utility heterogeneity may invalidate population-level conclusions.

stat.ME↗

The super learner for time-to-event outcomes: A tutorial

Estimating risks or survival probabilities conditional on individual characteristics based on censored time-to-event data is a commonly faced task. This may be for the purpose of developing a prediction model or may be part of a wider estimation procedure, such as in causal inference. A challenge is that it is impossible to know at the outset which of a set of candidate models will provide the best risk estimates. The super learner is a powerful approach for finding the best model or combination of models ('ensemble') among a pre-specified set of candidate models or 'learners', which can include both 'statistical' models (e.g. parametric, semi-parametric models) and 'machine learning' models. Super learners for time-to-event outcomes have been developed, but the literature is technical and the full details of how these methods work and can be implemented in practice have not previously been presented in an accessible format. In this paper we provide a practical tutorial on super learner methods for time-to-event outcomes. An overview of the general steps involved in the super learner is given, followed by details of three specific implementations for time-to-event outcomes. These include the originally proposed super learner, which involves using a discrete time scale, and two more recently proposed versions of the super learner for continuous-time data. We compare the properties of the methods and provide information on how they can be implemented in R. The methods are illustrated using an open access data set and R code is provided.

stat.ME↗

The risks of risk assessment: causal blind spots when using prediction models for treatment decisions

Clinicians increasingly rely on prediction models to guide treatment choices. Most prediction models, however, are developed using observational data that include some patients who have already received the treatment the prediction model is meant to inform. Special attention to the causal role of those earlier treatments is required when interpreting the resulting predictions. We identify 'causal blind spots' in three common approaches to handling treatment when developing a prediction model: including treatment as a predictor, restricting to individuals taking a certain treatment, and ignoring treatment. Through several real examples, we illustrate how the risks obtained from models developed using such approaches may be misinterpreted and can lead to misinformed decision-making. Our discussion covers issues attributable to confounding, selection, mediation and changes in treatment protocols over time. We advocate for an extension of guidelines for the development, reporting and evaluation of prediction models to avoid such misinterpretations. Developers must ensure that the intended target population for the model, and the treatment conditions under which predictions hold, are clearly communicated. When prediction models are intended to inform treatment decisions, they need to provide estimates of risk under the specific treatment (or intervention) options being considered, known as 'prediction under interventions'. Next to suitable data, this requires causal reasoning and causal inference techniques during model development and evaluation. Being clear about what a given prediction model can and cannot be used for prevents misinformed treatment decisions and thereby prevents potential harm to patients.

stat.ME↗

Doubly robust estimation of marginal cumulative incidence curves for competing risk analysis

Covariate imbalance between treatment groups makes it difficult to compare cumulative incidence curves in competing risk analyses. In this paper we discuss different methods to estimate adjusted cumulative incidence curves including inverse probability of treatment weighting and outcome regression modeling. For these methods to work, correct specification of the propensity score model or outcome regression model, respectively, is needed. We introduce a new doubly robust estimator, which requires correct specification of only one of the two models. We conduct a simulation study to assess the performance of these three methods, including scenarios with model misspecification of the relationship between covariates and treatment and/or outcome. We illustrate their usage in a cohort study of breast cancer patients estimating covariate-adjusted marginal cumulative incidence curves for recurrence, second primary tumour development and death after undergoing mastectomy treatment or breast-conserving therapy. Our study points out the advantages and disadvantages of each covariate adjustment method when applied in competing risk analysis.

stat.ME↗

When accurate prediction models yield harmful self-fulfilling prophecies

Prediction models are popular in medical research and practice. By predicting an outcome of interest for specific patients, these models may help inform difficult treatment decisions, and are often hailed as the poster children for personalized, data-driven healthcare. We show however, that using prediction models for decision making can lead to harmful decisions, even when the predictions exhibit good discrimination after deployment. These models are harmful self-fulfilling prophecies: their deployment harms a group of patients but the worse outcome of these patients does not invalidate the predictive power of the model. Our main result is a formal characterization of a set of such prediction models. Next we show that models that are well calibrated before and after deployment are useless for decision making as they made no change in the data distribution. These results point to the need to revise standard practices for validation, deployment and evaluation of prediction models that are used in medical decisions.

stat.ME↗

Time-lag bias induced by unobserved heterogeneity: comparing treated patients to controls with a different start of follow-up

In comparative effectiveness research, treated and control patients might have a different start of follow-up as treatment is often started later in the disease trajectory. This typically occurs when data from treated and controls are not collected within the same source. Only patients who did not yet experience the event of interest whilst in the control condition end up in the treatment data source. In case of unobserved heterogeneity, these treated patients will have a lower average risk than the controls. We illustrate how failing to account for this time-lag between treated and controls leads to bias in the estimated treatment effect. We define estimands and time axes, then explore five methods to adjust for this time-lag bias by utilising the time between diagnosis and treatment initiation in different ways. We conducted a simulation study to evaluate whether these methods reduce the bias and then applied the methods to a comparison between fertility patients treated with insemination and similar but untreated patients. We conclude that time-lag bias can be vast and that the time between diagnosis and treatment initiation should be taken into account in the analysis to respect the chronology of the disease and treatment trajectory.

stat.ME↗

Prediction under interventions: evaluation of counterfactual performance using longitudinal observational data

Predictions under interventions are estimates of what a person's risk of an outcome would be if they were to follow a particular treatment strategy, given their individual characteristics. Such predictions can give important input to medical decision making. However, evaluating predictive performance of interventional predictions is challenging. Standard ways of evaluating predictive performance do not apply when using observational data, because prediction under interventions involves obtaining predictions of the outcome under conditions that are different to those that are observed for a subset of individuals in the validation dataset. This work describes methods for evaluating counterfactual performance of predictions under interventions for time-to-event outcomes. This means we aim to assess how well predictions would match the validation data if all individuals had followed the treatment strategy under which predictions are made. We focus on counterfactual performance evaluation using longitudinal observational data, and under treatment strategies that involve sustaining a particular treatment regime over time. We introduce an estimation approach using artificial censoring and inverse probability weighting which involves creating a validation dataset that mimics the treatment strategy under which predictions are made. We extend measures of calibration, discrimination (c-index and cumulative/dynamic AUCt) and overall prediction error (Brier score) to allow assessment of counterfactual performance. The methods are evaluated using a simulation study, including scenarios in which the methods should detect poor performance. Applying our methods in the context of liver transplantation shows that our procedure allows quantification of the performance of predictions supporting crucial decisions on organ allocation.

stat.ME↗

Risk-based decision making: estimands for sequential prediction under interventions

Prediction models are used amongst others to inform medical decisions on interventions. Typically, individuals with high risks of adverse outcomes are advised to undergo an intervention while those at low risk are advised to refrain from it. Standard prediction models do not always provide risks that are relevant to inform such decisions: e.g., an individual may be estimated to be at low risk because similar individuals in the past received an intervention which lowered their risk. Therefore, prediction models supporting decisions should target risks belonging to defined intervention strategies. Previous works on prediction under interventions assumed that the prediction model was used only at one time point to make an intervention decision. In clinical practice, intervention decisions are rarely made only once: they might be repeated, deferred and re-evaluated. This requires estimated risks under interventions that can be reconsidered at several potential decision moments. In the current work, we highlight key considerations for formulating estimands in sequential prediction under interventions that can inform such intervention decisions. We illustrate these considerations by giving examples of estimands for a case study about choosing between vaginal delivery and cesarean section for women giving birth. Our formalization of prediction tasks in a sequential, causal, and estimand context provides guidance for future studies to ensure that the right question is answered and appropriate causal estimation approaches are chosen to develop sequential prediction models that can inform intervention decisions.

stat.ME↗

Prediction meets causal inference: the role of treatment in clinical prediction models

In this paper we study approaches for dealing with treatment when developing a clinical prediction model. Analogous to the estimand framework recently proposed by the European Medicines Agency for clinical trials, we propose a `predictimand' framework of different questions that may be of interest when predicting risk in relation to treatment started after baseline. We provide a formal definition of the estimands matching these questions, give examples of settings in which each is useful and discuss appropriate estimators including their assumptions. We illustrate the impact of the predictimand choice in a dataset of patients with end-stage kidney disease. We argue that clearly defining the estimand is equally important in prediction research as in causal inference.

stat.ME↗