arXiv · 2309.06723
PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction Network
Abstract
It is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. However, previous studies on audio-visual speaker extraction have not effectively handled the varying talking face. This paper studies how to take full advantage of the varying talking face. We propose a Pose-Invariant Audio-Visual Speaker Extraction Network (PIAVE) that incorporates an additional pose-invariant view to improve audio-visual speaker extraction. Specifically, we generate the pose-invariant view from each original pose orientation, which enables the model to receive a consistent frontal view of the talker regardless of his/her head pose, therefore, forming a multi-view visual input for the speaker. Experiments on the multi-view MEAD and in-the-wild LRS3 dataset demonstrate that PIAVE outperforms the state-of-the-art and is more robust to pose variations.
Explore related subjects
Keep this discovery
Qinghua Liu, Meng Ge, Zhizheng Wu, Haizhou Li. 2023-09-13. PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction Network. https://doi.org/10.21437/interspeech.2023-889
Cite the original work for its findings. Save a collection to share your selection of sources.