arXiv · 2609.32976
Adaptive Ensemble Selection for Noisy Labels on Tabular Data
Abstract
Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label noise is critical for building reliable models. We propose a data-centric reasoning module for AI data science systems that automatically diagnoses dataset quality and selects appropriate cleaning strategies. Given a dataset, a meta-model predicts weights over a diverse set of detectors, including confidence-based, neighborhood-based, and distributional methods. Across benchmark datasets with controlled noise, our approach achieves performance comparable to a Confident Learning baseline on average, with dataset-dependent gains and losses, particularly in heterogeneous regimes. We further show that detector effectiveness is systematically linked to dataset properties. These results demonstrate the value of descriptor-driven, data-centric ensembling as a component of AI-assisted data-science pipelines for robust dataset assessment and model reliability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Faizaan Ali, Inwon Kang, Oshani Seneviratne. 2026-09-26. Adaptive Ensemble Selection for Noisy Labels on Tabular Data. https://arxiv.org/abs/2609.32976
Cite the original work for its findings. Save a collection to share your selection of sources.