arXiv · 2205.07180
Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT
Abstract
This paper investigates self-supervised pre-training for audio-visual speaker representation learning where a visual stream showing the speaker's mouth area is used alongside speech as inputs. Our study focuses on the Audio-Visual Hidden Unit BERT (AV-HuBERT) approach, a recently developed general-purpose audio-visual speech pre-training framework. We conducted extensive experiments probing the effectiveness of pre-training and visual modality. Experimental results suggest that AV-HuBERT generalizes decently to speaker related downstream tasks, improving label efficiency by roughly ten fold for both audio-only and audio-visual speaker verification. We also show that incorporating visual information, even just the lip area, greatly improves the performance and noise robustness, reducing EER by 38% in the clean condition and 75% in noisy conditions.
Explore related subjects
Keep this discovery
Bowen Shi, Abdelrahman Mohamed, Wei-Ning Hsu. 2022-05-15. Learning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT. https://arxiv.org/abs/2205.07180
Cite the original work for its findings. Save a collection to share your selection of sources.