arXiv · 2408.06827
PRESENT: Zero-Shot Text-to-Prosody Control
Abstract
Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-to-speech (TTS) models, we present PRESENT (PRosody Editing without Style Embeddings or New Training), which exploits explicit prosody prediction in FastSpeech2-based models by modifying the inference process directly. We apply our text-to-prosody framework to zero-shot language transfer using a JETS model exclusively trained on English LJSpeech data. We obtain character error rates (CER) of 12.8%, 18.7% and 5.9% for German, Hungarian and Spanish respectively, beating the previous state-of-the-art CER by over 2x for all three languages. Furthermore, we allow subphoneme-level control, a first in this field. To evaluate its effectiveness, we show that PRESENT can improve the prosody of questions, and use it to generate Mandarin, a tonal language where vowel pitch varies at subphoneme level. We attain 25.3% hanzi CER and 13.0% pinyin CER with the JETS model. All our code and audio samples are available online.
Explore related subjects
Keep this discovery
Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman, Dorien Herremans. 2024-08-13. PRESENT: Zero-Shot Text-to-Prosody Control. https://arxiv.org/abs/2408.06827
Cite the original work for its findings. Save a collection to share your selection of sources.