arXiv · 2606.19668
Code-Switching Reveals Language Anchoring in Multilingual LLMs
Abstract
Multilingual Large Language Models (MLLMs) are increasingly expected to handle Code-Switched (CS) inputs, yet mixing languages frequently degrades performance relative to source- or target-language monolingual counterparts. To understand this degradation, we use grammar-forced CS as a controlled diagnostic setting for locating CS representations relative to their source and target counterparts. We introduce Anchor Bias, a geometric measure that quantifies language anchoring, whether a CS hidden state aligns closer to its source or target language counterpart. Across diverse MLLMs, Anchor Bias reveals a consistent grammar-frame effect: source-framed CS stays source-anchored, whereas target-framed CS shifts target-ward and shows larger Question Answering (QA) degradation. Motivated by this representational pattern, we propose CANVAS (Contextual Anchor-based Neural Vector Alignment Steering), an inference-time intervention that extracts a source-side canvas from the input and softly steers target-language hidden states toward the source anchor during prefill. CANVAS consistently recovers QA F1 across MLLMs and CS conditions, showing that internal anchoring signals provide an actionable target for mitigating CS inference failures.
Explore related subjects
Keep this discovery
Jeonghyun Park, Seunghyun Yoon, Yonghyun Jun, Hwanhee Lee. 2026-06-18. Code-Switching Reveals Language Anchoring in Multilingual LLMs. https://arxiv.org/abs/2606.19668
Cite the original work for its findings. Save a collection to share your selection of sources.