arXiv · 2608.21188
SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks
Abstract
Diffusion-based models are emerging in the speech enhancement domain and are achieving state-of-the-art performance across various benchmark datasets. A major downside of diffusion models is that data generation requires many evaluations of a typically large neural network, which results in high overall complexity. In this work, we propose a slimmable diffusion model that employs adaptive network widths throughout the data generation process to reduce computational cost. By using a greedy search algorithm to optimize the network width schedule, our method achieves performance comparable to baseline diffusion models with significantly reduced computational complexity. Notably, our approach reduces the computational complexity by up to $87.5\%$ without a significant drop in objective metrics, such as perceptual evaluation of speech quality (PESQ) and SI-SDR.
Explore related subjects
Keep this discovery
Nagashree K. S. Rao, Shrishti Saha Shetu, Mohamed Elminshawi, Emanuël A. P. Habets, Andreas Brendel. 2026-08-21. SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks. https://arxiv.org/abs/2608.21188
Cite the original work for its findings. Save a collection to share your selection of sources.