arXiv · 2203.15081
Word Discovery in Visually Grounded, Self-Supervised Speech Models
Abstract
We present a method for visually-grounded spoken term discovery. After training either a HuBERT or wav2vec2.0 model to associate spoken captions with natural images, we show that powerful word segmentation and clustering capability emerges within the model's self-attention heads. Our experiments reveal that this ability is not present to nearly the same extent in the base HuBERT and wav2vec2.0 models, suggesting that the visual grounding task is a crucial component of the word discovery capability we observe. We also evaluate our method on the Buckeye word segmentation and ZeroSpeech spoken term discovery tasks, where we perform on par with or better than currently published methods on several metrics. Code and model weights are available at https://github.com/jasonppy/word-discovery.
Explore related subjects
Keep this discovery
Puyuan Peng, David Harwath. 2022-03-28. Word Discovery in Visually Grounded, Self-Supervised Speech Models. https://arxiv.org/abs/2203.15081
Cite the original work for its findings. Save a collection to share your selection of sources.