arXiv · 2609.33734
Yorùbá in Unicode: An Overview of a Problem
Abstract
There is a recurrent problem in the writing of Yorùbá on the internet and on the computer that has proven intractable over the years. The language, along with other African languages that depend on diacritics for disambiguation, requires a small set of precomposed characters that Unicode does not encode. This has forced writers and digital systems to rely on combining character sequences that behave inconsistently across platforms, corrupt under font substitution, and fail in search. This paper documents that failure across a range of real world contexts, from published books to web platforms to mobile keyboards, using personal and empirical evidence. It identifies Unicode's NFC normalization stability policy as the structural constraint that prevents a straightforward fix, arguing for direct intervention of the Consortium in solving the active problem, proposing a formal encoding request for the four core Yorùbá characters as the most durable path to resolution.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kólá Túbòsún. 2026-09-27. Yorùbá in Unicode: An Overview of a Problem. https://arxiv.org/abs/2609.33734
Cite the original work for its findings. Save a collection to share your selection of sources.