arXiv · 2104.11985
Language ID Prediction from Speech Using Self-Attentive Pooling and 1D-Convolutions
Abstract
This memo describes NTR-TSU submission for SIGTYP 2021 Shared Task on predicting language IDs from speech. Spoken Language Identification (LID) is an important step in a multilingual Automated Speech Recognition (ASR) system pipeline. For many low-resource and endangered languages, only single-speaker recordings may be available, demanding a need for domain and speaker-invariant language ID systems. In this memo, we show that a convolutional neural network with a Self-Attentive Pooling layer shows promising results for the language identification task.
Explore related subjects
Keep this discovery
Roman Bedyakin, Nikolay Mikhaylovskiy. 2021-04-24. Language ID Prediction from Speech Using Self-Attentive Pooling and 1D-Convolutions. https://doi.org/10.18653/v1/2021.sigtyp-1.12
Cite the original work for its findings. Save a collection to share your selection of sources.