arXiv · 2312.11011
VinaLLaMA: LLaMA-based Vietnamese Foundation Model
Abstract
In this technical report, we present VinaLLaMA, an open-weight, state-of-the-art (SOTA) Large Language Model for the Vietnamese language, built upon LLaMA-2 with an additional 800 billion trained tokens. VinaLLaMA not only demonstrates fluency in Vietnamese but also exhibits a profound understanding of Vietnamese culture, making it a truly indigenous model. VinaLLaMA-7B-chat, trained on 1 million high-quality synthetic samples, achieves SOTA results on key benchmarks, including VLSP, VMLU, and Vicuna Benchmark Vietnamese, marking a significant advancement in the Vietnamese AI landscape and offering a versatile resource for various applications.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Quan Nguyen, Huy Pham, Dung Dao. 2023-12-18. VinaLLaMA: LLaMA-based Vietnamese Foundation Model. https://arxiv.org/abs/2312.11011
Cite the original work for its findings. Save a collection to share your selection of sources.