arXiv · 2306.15392
Assessing Dataset Quality Through Decision Tree Characteristics in Autoencoder-Processed Spaces
Abstract
In this paper, we delve into the critical aspect of dataset quality assessment in machine learning classification tasks. Leveraging a variety of nine distinct datasets, each crafted for classification tasks with varying complexity levels, we illustrate the profound impact of dataset quality on model training and performance. We further introduce two additional datasets designed to represent specific data conditions - one maximizing entropy and the other demonstrating high redundancy. Our findings underscore the importance of appropriate feature selection, adequate data volume, and data quality in achieving high-performing machine learning models. To aid researchers and practitioners, we propose a comprehensive framework for dataset quality assessment, which can help evaluate if the dataset at hand is sufficient and of the required quality for specific tasks. This research offers valuable insights into data assessment practices, contributing to the development of more accurate and robust machine learning models.
Explore related subjects
Keep this discovery
Szymon Mazurek, Maciej Wielgosz. 2023-06-27. Assessing Dataset Quality Through Decision Tree Characteristics in Autoencoder-Processed Spaces. https://arxiv.org/abs/2306.15392
Cite the original work for its findings. Save a collection to share your selection of sources.