arXiv · 2503.13921
Learning Over Dirty Data with Minimal Repairs
Abstract
Missing data often exists in real-world datasets, requiring significant time and effort for data repair to learn accurate models. In this paper, we show that imputing all missing values is not always necessary to achieve an accurate ML model. We introduce concepts of minimal and almost minimal repair, which are subsets of missing data items in training data whose imputation delivers accurate and reasonably accurate models, respectively. Imputing these subsets can significantly reduce the time, computational resources, and manual effort required for learning. We show that finding these subsets is NP-hard for some popular models and propose efficient approximation algorithms for wide range of models. Our extensive experiments indicate that our proposed algorithms can substantially reduce the time and effort required to learn on incomplete datasets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cheng Zhen, Prayoga, Nischal Aryal, Arash Termehchy, Garrett Biwer, Lubna Alzamil. 2025-03-18. Learning Over Dirty Data with Minimal Repairs. https://arxiv.org/abs/2503.13921
Cite the original work for its findings. Save a collection to share your selection of sources.