arXiv · 2412.06484
Small Languages, Big Models: A Study of Continual Training on Languages of Norway
Abstract
Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern S\'ami. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokm\r{a}l, Nynorsk, and Northern S\'ami with 11.4 billion parameters: NorMistral-11B.
Explore related subjects
Keep this discovery
David Samuel, Vladislav Mikhailov, Erik Velldal, Lilja Øvrelid, Lucas Georges Gabriel Charpentier, Andrey Kutuzov, Stephan Oepen. 2024-12-09. Small Languages, Big Models: A Study of Continual Training on Languages of Norway. https://arxiv.org/abs/2412.06484
Cite the original work for its findings. Save a collection to share your selection of sources.