arXiv · 2609.34223
Uncovering Ordinal-Matching Bias in Audio-Visual LLMs
Abstract
This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27\% for Qwen2.5-Omni and 2.57\% for video-SALMONN2+ across three audio-visual benchmarks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jihoo Jung, Youngjoon Jang, Hyebin Cho, Suho Yoo, Joon Son Chung. 2026-09-28. Uncovering Ordinal-Matching Bias in Audio-Visual LLMs. https://arxiv.org/abs/2609.34223
Cite the original work for its findings. Save a collection to share your selection of sources.