arXiv · 2008.07244
Efficient Low-Latency Speech Enhancement with Mobile Audio Streaming Networks
Abstract
We propose Mobile Audio Streaming Networks (MASnet) for efficient low-latency speech enhancement, which is particularly suitable for mobile devices and other applications where computational capacity is a limitation. MASnet processes linear-scale spectrograms, transforming successive noisy frames into complex-valued ratio masks which are then applied to the respective noisy frames. MASnet can operate in a low-latency incremental inference mode which matches the complexity of layer-by-layer batch mode. Compared to a similar fully-convolutional architecture, MASnet incorporates depthwise and pointwise convolutions for a large reduction in fused multiply-accumulate operations per second (FMA/s), at the cost of some reduction in SNR.
Explore related subjects
Keep this discovery
Michał Romaniuk, Piotr Masztalski, Karol Piaskowski, Mateusz Matuszewski. 2020-08-17. Efficient Low-Latency Speech Enhancement with Mobile Audio Streaming Networks. https://arxiv.org/abs/2008.07244
Cite the original work for its findings. Save a collection to share your selection of sources.