arXiv · 2505.19616
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
Abstract
Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet often exhibit poor robustness when exposed to spurious modality interference, such as irrelevant text in vision understanding, or irrelevant visual content in question answering. At its core, modality interference refers to cases where spurious signals from non-essential modalities distort model decisions, which we systematically analyze through causal, perturbation-based diagnostic experiments. To address this problem, we propose a unified finetuning framework that combines heuristic and adversarial perturbation-based data augmentation with output-level consistency regularization between original and perturbed inputs. Extensive experiments across image-heavy, text-heavy, and multimodal benchmarks, spanning multiple MLLM architectures and model scales, demonstrate consistent improvements in unimodal robustness and generalization, while improving standard multimodal performance.
Explore related subjects
Keep this discovery
Rui Cai, Bangzheng Li, Xiaofei Wen, Muhao Chen, Zhe Zhao. 2025-05-26. Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models. https://arxiv.org/abs/2505.19616
Cite the original work for its findings. Save a collection to share your selection of sources.