arXiv · 2606.18924
Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs
Abstract
While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models favor text over acoustic evidence, potentially leading to hallucinated responses. However, the internal mechanisms underlying how these models behave when audio and textual inputs contradict each other remain unexplored. In this work, we present the first mechanistic analysis of this phenomenon by tracing the propagation of internal representations across layers. Our investigation reveals three key findings: (i) text dominance is consistently observed across models; (ii) while text and audio rely on functionally distinct pathways, they ultimately converge into a shared semantic space in late layers; and (iii) the text pathway does not erase audio information, but rather actively suppresses intact audio representations. Building on these insights, we leverage back-patching, a training-free intervention that routes late-layer audio activations back into earlier layers. This amplifies the audio representations, enabling them to overcome textual suppression. Our evaluation shows that back-patching consistently reduces text dominance, demonstrating a mechanistic route to mitigating text dominance under conflict.
Explore related subjects
Keep this discovery
Hyebin Cho, Suho Yoo, Jaehyuk Jang, Changick Kim, Joon Son Chung. 2026-09-04. Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs. https://arxiv.org/abs/2606.18924
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.