arXiv · 2604.22887
StackFeat: a convergent algorithm for optimal predictor selection in genomic data
Abstract
In high-dimensional genomic data, the curse of dimensionality (d >> n) and limited sampling make feature selection inherently unstable - a critical barrier to biomarker discovery. We introduce StackFeat, an iterative algorithm that accumulates two statistics across repeated cross-validation: signed coefficients (measuring effect strength and direction) and selection frequencies (estimating selection probability). Only features ranking highly by both criteria are retained. On a COVID-19 miRNA dataset (GSE240888), StackFeat identified a stable 5-miRNA signature from 332 features (98.5% reduction), achieving AUC 0.922, significantly outperforming the benchmark 9-gene set (AUC 0.907, p = 0.0016). The signature includes hsa-miR-150-5p, a marker implicated in both COVID-19 survival and Dengue infection. This dual-criterion approach provides convergence guarantees absent in single-criterion methods, enabling discovery of known biomarkers, novel candidates, and previously unknown relationships. Keywords: marker selection, feature selection, bioinformatics, dimensionality reduction, robust algorithm, stacking, miRNA, COVID-19
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Akbar Yermekov, D. A. Herrera-Martí. 2026-04-24. StackFeat: a convergent algorithm for optimal predictor selection in genomic data. https://arxiv.org/abs/2604.22887
Cite the original work for its findings. Save a collection to share your selection of sources.