arXiv · 2603.14032
Beyond Two-stage Diffusion TTS: Joint Structure and Content Refinement via Jump Diffusion
Abstract
Diffusion and flow matching TTS faces a tension between discrete temporal structure and continuous spectral modeling. Two-stage models diffuse on fixed alignments, often collapsing to mean prosody; single-stage models avoid explicit durations but suffer alignment instability. We propose a jump-diffusion framework where discrete jumps model temporal structure and continuous diffusion refines spectral content within one process. Even in its one-shot degenerate form, our framework achieves 3.37% WER vs. 4.38% for Grad-TTS with improved UTMOSv2 on LJSpeech. The full iterative UDD variant further enables adaptive prosody, autonomously inserting natural pauses in out-of-distribution slow speech rather than stretching uniformly. Audio samples are available at https://anonymousinterpseech.github.io/TTS_Demo/.
Explore related subjects
Keep this discovery
Jiabao Ai, Minghui Zhao, Anton Ragni. 2026-03-14. Beyond Two-stage Diffusion TTS: Joint Structure and Content Refinement via Jump Diffusion. https://arxiv.org/abs/2603.14032
Cite the original work for its findings. Save a collection to share your selection of sources.