arXiv · 1602.01925
Massively Multilingual Word Embeddings
Abstract
We introduce new methods for estimating and evaluating embeddings of words in more than fifty languages in a single shared embedding space. Our estimation methods, multiCluster and multiCCA, use dictionaries and monolingual data; they do not require parallel data. Our new evaluation method, multiQVEC-CCA, is shown to correlate better than previous ones with two downstream tasks (text categorization and parsing). We also describe a web portal for evaluation that will facilitate further research in this area, along with open-source releases of all our methods.
Explore related subjects
Keep this discovery
Waleed Ammar, George Mulcaire, Yulia Tsvetkov, Guillaume Lample, Chris Dyer, Noah A. Smith. 2016-02-05. Massively Multilingual Word Embeddings. https://arxiv.org/abs/1602.01925
Cite the original work for its findings. Save a collection to share your selection of sources.