arXiv · 2412.20218
YAD: Leveraging T5 for Improved Automatic Diacritization of Yor\`ub\'a Text
Abstract
In this work, we present Yor\`ub\'a automatic diacritization (YAD) benchmark dataset for evaluating Yor\`ub\'a diacritization systems. In addition, we pre-train text-to-text transformer, T5 model for Yor\`ub\'a and showed that this model outperform several multilingually trained T5 models. Lastly, we showed that more data and larger models are better at diacritization for Yor\`ub\'a
Explore related subjects
Keep this discovery
Akindele Michael Olawole, Jesujoba O. Alabi, Aderonke Busayo Sakpere, David I. Adelani. 2024-12-28. YAD: Leveraging T5 for Improved Automatic Diacritization of Yor\`ub\'a Text. https://arxiv.org/abs/2412.20218
Cite the original work for its findings. Save a collection to share your selection of sources.