Searcharxiv⌕ Search

arXiv subjects

Ewout W Steyerberg

Publications and source records attributed to Ewout W Steyerberg.

3 recordsLinked to original sources

Instability in Patient Clustering: A Multiverse Analysis of Unsupervised Clustering in the CENTER-TBI cohort

Understanding patient heterogeneity is key to improving prognostic modeling in traumatic brain injury (TBI). Unsupervised clustering is widely used to explore patterns in patient characteristics that may define subgroups. However, it involves a multitude of decisions, including the choice of algorithm, the distance metric, and the method used to determine the "optimal" number of clusters. The aim of this study is to investigate how these choices influence the resulting clustering solution. We analyzed data from 4,509 patients enrolled in the Collaborative European NeuroTrauma Effectiveness Research in TBI (CENTER-TBI) study. K-medoids, agglomerative, and spectral clustering were applied in a complete 3 X 2 X 2 factorial design, in combination with Euclidean or Gower's distances, and silhouette score or gap statistic to choose the number of clusters. We investigated the agreement of clustering solutions with UpSet Plots and stability with the (adjusted) Rand index. Comparisons were made both across approaches using the original dataset and within approaches using bootstrap resampling. Clustering results varied substantially depending on the analysis choices. The number of suggested clusters varied widely, from one to twenty-five. Adjusted Rand indices confirmed low concordance between methods. Moreover, none of the clustering solutions demonstrated discriminatory performance comparable to a supervised logistic regression model in classifying patient recovery illustrating the limited usefulness of clustering for this purpose. The high instability in clustering results compromises interpretability and underscores that such solutions should not be blindly interpreted as underlying structure.

stat.ME↗

The fundamental problem of risk prediction for individuals: health AI, uncertainty, and personalized medicine

Background and Objective: Clinical prediction models are commonly evaluated regarding performance for a population, although decisions are made for individuals. The classic view relates uncertainty in risk estimates for individuals to sample size (estimation uncertainty) while other sources are model uncertainty (variability in modeling choices) and applicability uncertainty (variability in measurement procedures and between populations). We aim to illustrate the uncertainty of prediction models in estimating individual risks with an ovarian cancer example. Methods: We used real and synthetic data for ovarian cancer diagnosis to train 59400 models with variations in estimation, model, and applicability uncertainty. We then used these models to estimate the probability of ovarian cancer in a fixed test set of 100 patients and evaluate the variability in individual estimates. Results: We show empirically that estimation uncertainty can be strongly dominated by model uncertainty and applicability uncertainty, even for models that perform well at the population level. Estimation uncertainty decreased considerably with increasing training sample size, whereas model and applicability uncertainty remained large. Conclusion: Individual risk estimates are far more uncertain than often assumed. Model uncertainty and applicability uncertainty usually remain invisible when prediction models or algorithms are based on a single study. Predictive algorithms should inform, not dictate, care and support personalization through clinician-patient interaction.

stat.ME↗

Mining the contribution of intensive care clinical course to outcome after traumatic brain injury

Existing methods to characterise the evolving condition of traumatic brain injury (TBI) patients in the intensive care unit (ICU) do not capture the context necessary for individualising treatment. Here, we integrate all heterogenous data stored in medical records (1,166 pre-ICU and ICU variables) to model the individualised contribution of clinical course to six-month functional outcome on the Glasgow Outcome Scale - Extended (GOSE). On a prospective cohort (n=1,550, 65 centres) of TBI patients, we train recurrent neural network models to map a token-embedded time series representation of all variables (including missing values) to an ordinal GOSE prognosis every two hours. The full range of variables explains up to 52% (95% CI: 50%-54%) of the ordinal variance in functional outcome. Up to 91% (95% CI: 90%-91%) of this explanation is derived from pre-ICU and admission information (i.e., static variables). Information collected in the ICU (i.e., dynamic variables) increases explanation (by up to 5% [95% CI: 4%-6%]), though not enough to counter poorer overall performance in longer-stay (>5.75 days) patients. Highest-contributing variables include physician-based prognoses, CT features, and markers of neurological function. Whilst static information currently accounts for the majority of functional outcome explanation after TBI, data-driven analysis highlights investigative avenues to improve dynamic characterisation of longer-stay patients. Moreover, our modelling strategy proves useful for converting large patient records into interpretable time series with missing data integration and minimal processing.

cs.LG↗