arXiv · 2501.11311
A2SB: Audio-to-Audio Schrodinger Bridges
Abstract
Real-world audio is often degraded by numerous factors. This work presents an audio restoration model tailored for high-res music at 44.1kHz. Our model, Audio-to-Audio Schr\"odinger Bridges (A2SB), is capable of both bandwidth extension (predicting high-frequency components) and inpainting (re-generating missing segments). Critically, A2SB is end-to-end requiring no vocoder to predict waveform outputs, able to restore hour-long audio inputs, and trained on permissively licensed music data. A2SB is capable of achieving state-of-the-art band-width extension and inpainting quality on several out-of-distribution music test sets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhifeng Kong, Kevin J Shih, Weili Nie, Arash Vahdat, Sang-gil Lee, Joao Felipe Santos, Ante Jukic, Rafael Valle, Bryan Catanzaro. 2025-01-20. A2SB: Audio-to-Audio Schrodinger Bridges. https://arxiv.org/abs/2501.11311
Cite the original work for its findings. Save a collection to share your selection of sources.