arXiv · 2601.09716
Opportunities and Challenges of Natural Language Processing for Low-Resource Senegalese Languages in Social Science Research
Abstract
Natural Language Processing (NLP) is rapidly transforming research methodologies across disciplines, yet African languages remain largely underrepresented in this technological shift. This paper provides the first comprehensive overview of NLP progress and challenges for the six national languages officially recognized by the Senegalese Constitution: Wolof, Pulaar, S\'er\`ere, Diola, Mandingue, and Sonink\'e. We synthesize linguistic, socio-technical, and infrastructural factors that shape their digital readiness and identify gaps in data, tools, and benchmarks. Building on existing initiatives and research works, we analyze ongoing efforts in various tasks, covering both text and speech modalities. We also provide a centralized GitHub repository that compiles publicly accessible resources for a range of NLP tasks across these languages, designed to facilitate collaboration and reproducibility. A special focus is devoted to the application of NLP to the social sciences, where multilingual transcription, translation, and retrieval pipelines can significantly enhance the efficiency and inclusiveness of field research. The paper concludes by outlining a roadmap toward sustainable, community-centered NLP ecosystems for Senegalese languages, emphasizing ethical data governance, open resources, and interdisciplinary collaboration.
Explore related subjects
Keep this discovery
Derguene Mbaye, Tatiana D. P. Mbengue, Madoune R. Seye, Moussa Diallo, Mamadou L. Ndiaye, Dimitri S. Adjanohoun, Cheikh S. Wade, Djiby Sow, Jean-Claude B. Munyaka, Jerome Chenal. 2025-12-24. Opportunities and Challenges of Natural Language Processing for Low-Resource Senegalese Languages in Social Science Research. https://arxiv.org/abs/2601.09716
Cite the original work for its findings. Save a collection to share your selection of sources.