arXiv · 2608.09941
The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs
Abstract
While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the quantization tax-remain overwhelmingly English-centric. We present a zero-shot multilingual evaluation of 4-bit quantization across the Gemma 4 and Qwen 3.5 architectures. Evaluating on eight typo-logically diverse languages using MMLU ProX Lite and GlobalPIQA, we show parameter truncation exposes deep pre-training inequalities. We identify four phenomena: (1) Typological Fragility: low-resource and specific non-Latin scripts suffer representational collapse via architecture-specific double dissociations, failing to generate valid task logits; (2) Home Language Fragility Paradox: foundational pre-training pathways provide limited precision loss protection; (3) Domain-Specific Forgetting: multi-step cross-lingual routing degrades while associative soft-science recall remains robust; and (4) Quantization Resistance: highly saturated, typologically aligned domains resist deterministic degradation, with post-quantization performance gains bounded by statistical noise.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mohammad Wathiq Soualhi. 2026-06-21. The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs. https://arxiv.org/abs/2608.09941
Cite the original work for its findings. Save a collection to share your selection of sources.