arXiv · 2202.12163
Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech
Abstract
In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry information in long-form audio via a recurrent form, such that the inference can be performed in a streaming fashion. Additionally, we investigate two domain adaptation approaches to allow adapting an existing language identification model without retraining the model parameters for a new domain. We perform a comparative study of different model topologies under different constraints of model size, and find that conformer-based models significantly outperform LSTM and transformer based models. Our experiments also show that attentive temporal pooling and domain adaptation improve model accuracy.
Explore related subjects
Keep this discovery
Quan Wang, Yang Yu, Jason Pelecanos, Yiling Huang, Ignacio Lopez Moreno. 2022-02-24. Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech. https://arxiv.org/abs/2202.12163
Cite the original work for its findings. Save a collection to share your selection of sources.