arXiv · 2608.17088
There is No Theoretical Curse of Multilinguality For Embedding Space Structure
Abstract
A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model. The curse of multilinguality describes the phenomenon of degradation in multilingual model performance as we increase language coverage, posing a threat to the above goal. This paper asks whether multilingual embedding spaces are inherently incapable of achieving perfect multilinguality without a prohibitive increase in required capacity. We first formalize the goal of "perfect multilinguality", embodied in two multilinguality conditions. We then prove that the minimum dimensionality required for perfect multilinguality grows only logarithmically in the number of languages. That is, we show that there is no theoretical curse of multilinguality for embedding space structure. This suggests that the empirical curse of multilinguality is a result of real world data and training conditions. We back this understanding with a small-scale empirical study. Our paper provides the first theoretical and intrinsic perspective on the curse of multilinguality.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Niyati Bafna, Neha Verma, Vilém Zouhar, Philipp Koehn, David Yarowsky. 2026-08-17. There is No Theoretical Curse of Multilinguality For Embedding Space Structure. https://arxiv.org/abs/2608.17088
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.