arXiv · 2102.06282
A reproduction of Apple's bi-directional LSTM models for language identification in short strings
Abstract
Language Identification is the task of identifying a document's language. For applications like automatic spell checker selection, language identification must use very short strings such as text message fragments. In this work, we reproduce a language identification architecture that Apple briefly sketched in a blog post. We confirm the bi-LSTM model's performance and find that it outperforms current open-source language identifiers. We further find that its language identification mistakes are due to confusion between related languages.
Explore related subjects
Keep this discovery
Mads Toftrup, Søren Asger Sørensen, Manuel R. Ciosici, Ira Assent. 2021-02-11. A reproduction of Apple's bi-directional LSTM models for language identification in short strings. https://arxiv.org/abs/2102.06282
Cite the original work for its findings. Save a collection to share your selection of sources.