arXiv · 2607.21814
Another look at predicting molecular breast cancer subtypes from the METABRIC data
Abstract
Classifying patients into different clusters based on genomic data can offer valuable insights into their projected disease-specific survival trajectories over time. Here we apply two supervised learning methods---Nearest Shrunken Centroids and LASSO--- to the METABRIC Breast Cancer data {metabric} of 1980 patients and 754 genes to perform this task. The {pamr} R package implements the Nearest Shrunken Centroids classifier and the { glmnet} R package is used to fit an ungrouped multinomial model, a grouped multinomial model, and a One-Versus-Rest model. Splitting our data into discovery and validation sets, we evaluate all four models' classification performance and the survival implications of their class predictions using cross validation and Kaplan-Meier curves. We find that the One-Versus-Rest model produces the lowest misclassification error of 0.0572 and the lowest median log-rank test of 0.380 statistic measuring the similarity between its Kaplan-Meier curves and the discovery set's true Kaplan-Meier curves. We show that a multinomial classification task split into several LASSO binomial classifiers offers promising results for patient clustering.
Explore related subjects
Keep this discovery
Isabella Lazarov, Robert Tibshirani. 2026-07-23. Another look at predicting molecular breast cancer subtypes from the METABRIC data. https://arxiv.org/abs/2607.21814
Cite the original work for its findings. Save a collection to share your selection of sources.