arXiv · 1803.09405
Automatic Identification of Closely-related Indian Languages: Resources and Experiments
Abstract
In this paper, we discuss an attempt to develop an automatic language identification system for 5 closely-related Indo-Aryan languages of India, Awadhi, Bhojpuri, Braj, Hindi and Magahi. We have compiled a comparable corpora of varying length for these languages from various resources. We discuss the method of creation of these corpora in detail. Using these corpora, a language identification system was developed, which currently gives state of the art accuracy of 96.48\%. We also used these corpora to study the similarity between the 5 languages at the lexical level, which is the first data-based study of the extent of closeness of these languages.
Explore related subjects
Keep this discovery
Ritesh Kumar, Bornini Lahiri, Deepak Alok, Atul Kr. Ojha, Mayank Jain, Abdul Basit, Yogesh Dawer. 2018-03-26. Automatic Identification of Closely-related Indian Languages: Resources and Experiments. https://arxiv.org/abs/1803.09405
Cite the original work for its findings. Save a collection to share your selection of sources.