arXiv · 2503.06547
Kr\'eyoLID From Language Identification Towards Language Mining
Abstract
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may be more appropriate to consider it a data mining problem. For these varieties, one knows ahead of time that the vast majority of documents are of little interest. By minimizing resources spent on classifying such documents, we can create corpora much faster and with better coverage than using established pipelines. To demonstrate the effectiveness of the language mining perspective, we introduce a new pipeline and corpora for several French-based Creoles.
Explore related subjects
Keep this discovery
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice, Benoît Sagot. 2025-03-09. Kr\'eyoLID From Language Identification Towards Language Mining. https://arxiv.org/abs/2503.06547
Cite the original work for its findings. Save a collection to share your selection of sources.