arXiv · 2510.04139
Fine Tuning Methods for Low-resource Languages
Abstract
The rise of Large Language Models has not been inclusive of all cultures. The models are mostly trained on English texts and culture which makes them underperform in other languages and cultural contexts. By developing a generalizable method for preparing culturally relevant datasets and post-training the Gemma 2 model, this project aimed to increase the performance of Gemma 2 for an underrepresented language and showcase how others can do the same to unlock the power of Generative AI in their country and preserve their cultural heritage.
Explore related subjects
Keep this discovery
Tim Bakkenes, Daniel Wang, Anton Johansson. 2025-10-05. Fine Tuning Methods for Low-resource Languages. https://arxiv.org/abs/2510.04139
Cite the original work for its findings. Save a collection to share your selection of sources.