arXiv · 2306.15636
On the Usefulness of Synthetic Tabular Data Generation
Abstract
Despite recent advances in synthetic data generation, the scientific community still lacks a unified consensus on its usefulness. It is commonly believed that synthetic data can be used for both data exchange and boosting machine learning (ML) training. Privacy-preserving synthetic data generation can accelerate data exchange for downstream tasks, but there is not enough evidence to show how or why synthetic data can boost ML training. In this study, we benchmarked ML performance using synthetic tabular data for four use cases: data sharing, data augmentation, class balancing, and data summarization. We observed marginal improvements for the balancing use case on some datasets. However, we conclude that there is not enough evidence to claim that synthetic tabular data is useful for ML training.
Explore related subjects
Keep this discovery
Dionysis Manousakas, Sergül Aydöre. 2023-06-27. On the Usefulness of Synthetic Tabular Data Generation. https://arxiv.org/abs/2306.15636
Cite the original work for its findings. Save a collection to share your selection of sources.