arXiv · 2312.14069
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
Abstract
We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis. We apply this to two tasks: speech resynthesis and speech-to-speech translation. In both cases, the benchmark evaluates the ability of the model to encode emphasis in the speech input and accurately reproduce it in the output, potentially across a change of speaker and language. As part of the evaluation pipeline, we introduce EmphaClass, a new model that classifies emphasis at the frame or word level.
Explore related subjects
Keep this discovery
Maureen de Seyssel, Antony D'Avirro, Adina Williams, Emmanuel Dupoux. 2023-12-21. EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models. https://arxiv.org/abs/2312.14069
Cite the original work for its findings. Save a collection to share your selection of sources.