arXiv · 2606.20001
Time-Unconditional Generative Speech Enhancement via Autonomous Rectified Flow
Abstract
Most generative speech enhancement methods rely on explicit time-step embeddings for temporal conditioning. In this paper, we propose the Autonomous Rectified Flow framework, which challenges the necessity of such conditioning. Using a linear interpolation path, we show that the target vector field is inherently time-invariant. We further introduce a time-unconditional network that eliminates explicit time-step information and infers the denoising direction solely from the spatial relationship between the current state and the noisy observation. Predicting this target vector field is equivalent to modeling the noise distribution. By avoiding overfitting to temporal trajectories, the proposed autonomous design significantly improves generation quality, robustness, and inference efficiency.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wen Zhang, Wenbin Jiang, Yang Zhang, Xiaofei Zhou. 2026-06-18. Time-Unconditional Generative Speech Enhancement via Autonomous Rectified Flow. https://arxiv.org/abs/2606.20001
Cite the original work for its findings. Save a collection to share your selection of sources.