arXiv · 2606.18485
MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data
Abstract
Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence boundary artifacts. Existing approaches either compress sequences, increase context length or naively concatenate independently synthesized chunks. We present an inference-time approach called MagpieTTS-LF that enables MagpieTTS to produce coherent long-form speech without model retraining. Our method introduces three key innovations: (1) soft attention priors to guide monotonic alignment while preserving past and future context; (2) a stateful inference algorithm that maintains context across sentence chunks, ensuring prosodic continuity; (3) history-aware text encoding that uses past text for discourse-level prosodic planning. Experiments on long texts show significant improvements in long-range intelligibility, prosodic coherence, speaker consistency, and boundary naturalness compared to other baselines.
Explore related subjects
Keep this discovery
Subhankar Ghosh, Jason Li, Paarth Neekhara, Shehzeen Hussain, Ryan Langman, Xuesong Yang, Roy Fejgin. 2026-06-16. MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data. https://arxiv.org/abs/2606.18485
Cite the original work for its findings. Save a collection to share your selection of sources.