arXiv · 2002.02848
Unsupervised pretraining transfers well across languages
Abstract
Cross-lingual and multi-lingual training of Automatic Speech Recognition (ASR) has been extensively investigated in the supervised setting. This assumes the existence of a parallel corpus of speech and orthographic transcriptions. Recently, contrastive predictive coding (CPC) algorithms have been proposed to pretrain ASR systems with unlabelled data. In this work, we investigate whether unsupervised pretraining transfers well across languages. We show that a slight modification of the CPC pretraining extracts features that transfer well to other languages, being on par or even outperforming supervised pretraining. This shows the potential of unsupervised methods for languages with few linguistic resources.
Explore related subjects
Keep this discovery
Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, Emmanuel Dupoux. 2020-02-07. Unsupervised pretraining transfers well across languages. https://arxiv.org/abs/2002.02848
Cite the original work for its findings. Save a collection to share your selection of sources.