arXiv · 2304.04157
An investigation of phrase break prediction in an End-to-End TTS system
Abstract
Purpose: This work explores the use of external phrase break prediction models to enhance listener comprehension in End-to-End Text-to-Speech (TTS) systems. Methods: The effectiveness of these models is evaluated based on listener preferences in subjective tests. Two approaches are explored: (1) a bidirectional LSTM model with task-specific embeddings trained from scratch, and (2) a pre-trained BERT model fine-tuned on phrase break prediction. Both models are trained on a multi-speaker English corpus to predict phrase break locations in text. The End-to-End TTS system used comprises a Tacotron2 model with Dynamic Convolutional Attention for mel spectrogram prediction and a WaveRNN vocoder for waveform generation. Results: The listening tests show a clear preference for text synthesized with predicted phrase breaks over text synthesized without them. Conclusion: These results confirm the value of incorporating external phrasing models within End-to-End TTS to enhance listener comprehension.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Anandaswarup Vadapalli. 2023-04-09. An investigation of phrase break prediction in an End-to-End TTS system. https://doi.org/10.1007/s42979-024-03652-0
Cite the original work for its findings. Save a collection to share your selection of sources.