arXiv · 2108.09814
UzBERT: pretraining a BERT model for Uzbek
Abstract
Pretrained language models based on the Transformer architecture have achieved state-of-the-art results in various natural language processing tasks such as part-of-speech tagging, named entity recognition, and question answering. However, no such monolingual model for the Uzbek language is publicly available. In this paper, we introduce UzBERT, a pretrained Uzbek language model based on the BERT architecture. Our model greatly outperforms multilingual BERT on masked language model accuracy. We make the model publicly available under the MIT open-source license.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
B. Mansurov, A. Mansurov. 2021-08-22. UzBERT: pretraining a BERT model for Uzbek. https://arxiv.org/abs/2108.09814
Cite the original work for its findings. Save a collection to share your selection of sources.