arXiv · 2507.15339
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
Abstract
Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments. Small models offer a potential alternative to large LLMs, yet still demand considerable data and compute. We present LionGuard 2, a lightweight, multilingual moderation classifier tailored to the Singapore context, supporting English, Chinese, Malay, and partial Tamil. Built on pre-trained OpenAI embeddings and a multi-head ordinal classifier, LionGuard 2 outperforms several commercial and open-source systems across 17 benchmarks, including both Singapore-specific and public English datasets. The system is actively deployed within the Singapore Government, demonstrating practical efficacy at scale. Our findings show that high-quality local data and robust multilingual embeddings can achieve strong moderation performance, without fine-tuning large models. We release our model weights and part of our training data to support future work on LLM safety.
Explore related subjects
Keep this discovery
Leanne Tan, Gabriel Chua, Ziyu Ge, Roy Ka-Wei Lee. 2025-07-21. LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators. https://arxiv.org/abs/2507.15339
Cite the original work for its findings. Save a collection to share your selection of sources.