arXiv · 2509.04032
What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages
Abstract
How similar are model outputs across languages? In this work, we study this question using a recently proposed model similarity metric $\kappa_p$ applied to 20 languages and 47 subjects in GlobalMMLU. Our analysis reveals that a model's responses become increasingly consistent across languages as its size and capability grow. Interestingly, models exhibit greater cross-lingual consistency within themselves than agreement with other models prompted in the same language. These results highlight not only the value of $\kappa_p$ as a practical tool for evaluating multilingual reliability, but also its potential to guide the development of more consistent multilingual systems.
Explore related subjects
Keep this discovery
Debangan Mishra, Arihant Rastogi, Agyeya Negi, Shashwat Goel, Ponnurangam Kumaraguru. 2025-09-04. What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages. https://arxiv.org/abs/2509.04032
Cite the original work for its findings. Save a collection to share your selection of sources.