arXiv · 2609.30984
THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer
Abstract
Text-to-speech needs written text in spoken form, and speech recognition output needs the reverse. For Khmer, neither direction has a maintained open-source tool, and the script makes both harder: words are not separated by spaces, and number words occur inside ordinary words. We present Tha, a Khmer text normalization and inverse text normalization toolkit built from weighted finite-state transducers. It segments and classifies a whole line in one shortest-path search, and a second transducer rejects token boundaries inside a Khmer syllable. On Google's Khmer test suite, Tha agrees with the reference on all 274 cardinals up to one spelling variant, and on 2,906 real TTS prompts, 153 of the 158 sentences it rewrites are correct. Tha is open source under the Apache 2.0 license.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Seanghay Yath. 2026-09-25. THA: Weighted Finite-State Text Normalization and Inverse Text Normalization for Khmer. https://arxiv.org/abs/2609.30984
Cite the original work for its findings. Save a collection to share your selection of sources.