arXiv · 2603.18425
Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs
Abstract
Task interference, the performance degradation caused by task switches within a single conversation, has been studied exclusively in text-only settings despite the growing prevalence of multimodal dialogue systems. We introduce a benchmark for evaluating this phenomenon in multimodal LLMs, covering six tasks across text and vision with systematic variation of history-target along three axes: modality mismatch, reasoning mismatch, and answer format mismatch. Experiments on both open-weights and proprietary models reveal that task interference is highly directional: switching from text-only to image-based targets causes severe performance drops, while the reverse transition yields minimal degradation. Interference is further amplified when mismatches co-occur across multiple dimensions, and is driven most strongly by modality differences, followed by answer format, while reasoning requirement shifts cause minimal degradation.
Explore related subjects
Keep this discovery
Masayuki Kawarada, Tatsuya Ishigaki, Hiroya Takamura. 2026-03-19. Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs. https://doi.org/10.63317/36ae8bm4re6t
Cite the original work for its findings. Save a collection to share your selection of sources.