arXiv · 2609.10392
Teacher-Free Self-Distilled Consistency Trajectory Learning for Fast Speech Enhancement
Abstract
Consistency trajectory models offer a route to fast, high-quality speech enhancement, collapsing the many reverse steps of diffusion-based enhancers into a handful. When instantiated on a Schr\"odinger bridge (SB), which pins the generative process to fixed clean and noisy endpoints, existing consistency-trajectory enhancers (SBCTMs) still require a pretrained teacher to supply trajectory supervision, which raises training cost and ties the final quality to that of the teacher. We propose a teacher-free, self-distilled consistency-trajectory framework that removes the external teacher: trajectory targets are generated by an exponential-moving-average (EMA) copy of the student, and the model is trained with a three-stage curriculum of $\x_0$ prediction, a self-distilled shortcut objective, and perceptual fine-tuning with a multi-resolution short-time Fourier transform (MR-STFT) loss. Using the same NCSN++ backbone as SBCTM, our model attains a wide-band PESQ of $3.01$, ESTOI $0.87$, and SI-SDR $19.07$\,dB on VoiceBank+DEMAND without a teacher. Varying step count and inference schedule we find that a geometric schedule at low reverse step count maximizes perceptual quality, while a higher-step uniform schedule favors signal fidelity, with the geometric advantage narrowing with reverse step count.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shuubham Ojha, Carol Espy-Wilson. 2026-09-09. Teacher-Free Self-Distilled Consistency Trajectory Learning for Fast Speech Enhancement. https://arxiv.org/abs/2609.10392
Cite the original work for its findings. Save a collection to share your selection of sources.