arXiv · 2310.09636
Generative Adversarial Training for Text-to-Speech Synthesis Based on Raw Phonetic Input and Explicit Prosody Modelling
Abstract
We describe an end-to-end speech synthesis system that uses generative adversarial training. We train our Vocoder for raw phoneme-to-audio conversion, using explicit phonetic, pitch and duration modeling. We experiment with several pre-trained models for contextualized and decontextualized word embeddings and we introduce a new method for highly expressive character voice matching, based on discreet style tokens.
Explore related subjects
Keep this discovery
Tiberiu Boros, Stefan Daniel Dumitrescu, Ionut Mironica, Radu Chivereanu. 2023-10-14. Generative Adversarial Training for Text-to-Speech Synthesis Based on Raw Phonetic Input and Explicit Prosody Modelling. https://arxiv.org/abs/2310.09636
Cite the original work for its findings. Save a collection to share your selection of sources.