arXiv · 2609.35397
Adapting offline models to a streaming context for music source separation
Abstract
Real-time music source separation must satisfy two constraints: a bound on algorithmic latency and a bound on computational cost. Offline separators are usually omitted from real-time comparisons or credited with a latency equal to their full input length. We show that this latency is set by where the output is read, not by the length of the separator's input. An unmodified offline model can therefore run in a streaming setting, without retraining. At each step, the input slides by one STFT hop, and one output hop is read out. The resulting latency can be as low as one STFT hop (23 ms), and the computational cost does not increase as latency shrinks. We identify a theoretical model-dependent latency boundary below which separation quality should drop steeply, and confirm this experimentally across three architectures. At equal algorithmic latency, streamed off-the-shelf checkpoints for HT-Demucs and SCNet match the published results of dedicated real-time models in terms of separation quality. Streamed models remain far less computationally efficient: only HT-Demucs runs faster than real time on a GPU.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dylan Sechet, Marc Evrard, Matthieu Kowalski. 2026-09-28. Adapting offline models to a streaming context for music source separation. https://arxiv.org/abs/2609.35397
Cite the original work for its findings. Save a collection to share your selection of sources.