arXiv · 2501.10811
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
Abstract
The technology for generating music from textual descriptions has seen rapid advancements. However, evaluating text-to-music (TTM) systems remains a significant challenge, primarily due to the difficulty of balancing performance and cost with existing objective and subjective evaluation methods. In this paper, we propose an automatic assessment task for TTM models to align with human perception. To address the TTM evaluation challenges posed by the professional requirements of music evaluation and the complexity of the relationship between text and music, we collect MusicEval, the first generative music assessment dataset. This dataset contains 2,748 music clips generated by 31 advanced and widely used models in response to 384 text prompts, along with 13,740 ratings from 14 music experts. Furthermore, we design a CLAP-based assessment model built on this dataset, and our experimental results validate the feasibility of the proposed task, providing a valuable reference for future development in TTM evaluation. The dataset is available at https://www.aishelltech.com/AISHELL_7A.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cheng Liu, Hui Wang, Jinghua Zhao, Shiwan Zhao, Hui Bu, Xin Xu, Jiaming Zhou, Haoqin Sun, Yong Qin. 2025-01-18. MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation. https://arxiv.org/abs/2501.10811
Cite the original work for its findings. Save a collection to share your selection of sources.