arXiv · 2507.22612
Adaptive Duration Model for Text Speech Alignment
Abstract
Speech-to-text alignment is a critical component of neural text to speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line, while non-autoregressive end to end TTS models rely on durations extracted from external sources. In this paper, we propose a novel duration prediction framework that can give promising phoneme-level duration distribution with given text. In our experiments, the proposed duration model has more precise prediction and adaptation ability to conditions, compared to previous baseline models. Specifically, it makes a considerable improvement on phoneme-level alignment accuracy and makes the performance of zero-shot TTS models more robust to the mismatch between prompt audio and input audio.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junjie Cao. 2025-07-30. Adaptive Duration Model for Text Speech Alignment. https://arxiv.org/abs/2507.22612
Cite the original work for its findings. Save a collection to share your selection of sources.