arXiv · 1911.05636
Prevalence of code mixing in semi-formal patient communication in low resource languages of South Africa
Abstract
In this paper we address the problem of code-mixing in resource-poor language settings. We examine data consisting of 182k unique questions generated by users of the MomConnect helpdesk, part of a national scale public health platform in South Africa. We show evidence of code-switching at the level of approximately 10% within this dataset -- a level that is likely to pose challenges for future services. We use a natural language processing library (Polyglot) that supports detection of 196 languages and attempt to evaluate its performance at identifying English, isiZulu and code-mixed questions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Monika Obrocka, Charles Copley, Themba Gqaza, Eli Grant. 2019-12-10. Prevalence of code mixing in semi-formal patient communication in low resource languages of South Africa. https://arxiv.org/abs/1911.05636
Cite the original work for its findings. Save a collection to share your selection of sources.