arXiv · 2206.06184
AmbiSep: Ambisonic-to-Ambisonic Reverberant Speech Separation Using Transformer Networks
Abstract
Consider a multichannel Ambisonic recording containing a mixture of several reverberant speech signals. Retreiving the reverberant Ambisonic signals corresponding to the individual speech sources blindly from the mixture is a challenging task as it requires to estimate multiple signal channels for each source. In this work, we propose AmbiSep, a deep neural network-based plane-wave domain masking approach to solve this task. The masking network uses learned feature representations and transformers in a triple-path processing configuration. We train and evaluate the proposed network architecture on a spatialized WSJ0-2mix dataset, and show that the method achieves a multichannel scale-invariant signal-to-distortion ratio improvement of 17.7 dB on the blind test set, while preserving the spatial characteristics of the separated sounds.
Explore related subjects
Keep this discovery
Adrian Herzog, Srikanth Raj Chetupalli, Emanuël A. P. Habets. 2022-06-13. AmbiSep: Ambisonic-to-Ambisonic Reverberant Speech Separation Using Transformer Networks. https://arxiv.org/abs/2206.06184
Cite the original work for its findings. Save a collection to share your selection of sources.