arXiv · 1811.07030
Exploring Tradeoffs in Models for Low-latency Speech Enhancement
Abstract
We explore a variety of neural networks configurations for one- and two-channel spectrogram-mask-based speech enhancement. Our best model improves on previous state-of-the-art performance on the CHiME2 speech enhancement task by 0.4 decibels in signal-to-distortion ratio (SDR). We examine trade-offs such as non-causal look-ahead, computation, and parameter count versus enhancement performance and find that zero-look-ahead models can achieve, on average, within 0.03 dB SDR of our best bidirectional model. Further, we find that 200 milliseconds of look-ahead is sufficient to achieve equivalent performance to our best bidirectional model.
Explore related subjects
Keep this discovery
Kevin Wilson, Michael Chinen, Jeremy Thorpe, Brian Patton, John Hershey, Rif A. Saurous, Jan Skoglund, Richard F. Lyon. 2018-11-16. Exploring Tradeoffs in Models for Low-latency Speech Enhancement. https://arxiv.org/abs/1811.07030
Cite the original work for its findings. Save a collection to share your selection of sources.