arXiv · 1606.02542
Symbolic Music Data Version 1.0
Abstract
In this document, we introduce a new dataset designed for training machine learning models of symbolic music data. Five datasets are provided, one of which is from a newly collected corpus of 20K midi files. We describe our preprocessing and cleaning pipeline, which includes the exclusion of a number of files based on scores from a previously developed probabilistic machine learning model. We also define training, testing and validation splits for the new dataset, based on a clustering scheme which we also describe. Some simple histograms are included.
Explore related subjects
Keep this discovery
Christian Walder. 2016-06-08. Symbolic Music Data Version 1.0. https://arxiv.org/abs/1606.02542
Cite the original work for its findings. Save a collection to share your selection of sources.