arXiv · 2412.14527
Statistical Undersampling with Mutual Information and Support Points
Abstract
Class imbalance and distributional differences in large datasets present significant challenges for classification tasks machine learning, often leading to biased models and poor predictive performance for minority classes. This work introduces two novel undersampling approaches: mutual information-based stratified simple random sampling and support points optimization. These methods prioritize representative data selection, effectively minimizing information loss. Empirical results across multiple classification tasks demonstrate that our methods outperform traditional undersampling techniques, achieving higher balanced classification accuracy. These findings highlight the potential of combining statistical concepts with machine learning to address class imbalance in practical applications.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alex Mak, Shubham Sahoo, Shivani Pandey, Yidan Yue, Linglong Kong. 2024-12-19. Statistical Undersampling with Mutual Information and Support Points. https://arxiv.org/abs/2412.14527
Cite the original work for its findings. Save a collection to share your selection of sources.