arXiv · 2607.04683
Failing to See or Failing to Know? Attributing Errors in Vision-Language Models
Abstract
Vision-language models (VLMs) can recognize entities in clear images yet still fail when answering questions that require factual knowledge beyond what is directly observable. Prior work has either examined individual failure modes in isolation or treated incorrect answers as monolithic, binary failures. We propose a tree-structured framework that organizes failures in knowledge-intensive visual question answering into model-specific operational outcomes. Across two datasets and four VLMs, we observe consistent distributions of operational outcomes: some failures occur before entity recognition, while others persist after the relevant entity is recognized. Visual token representations are most informative for recognition-related decisions. Prompt hidden states predict answer success more effectively, although factual-access attribution remains difficult and exhibits only a weak signal. These pre-generation signals support attribution-guided routing to targeted interventions, including image repair, entity support, question rewriting, and factual evidence.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Khang Nhat Hoang Vo, Artem Vazhentsev, Artem Shelmanov, Timothy Baldwin, Yova Kementchedjhieva. 2026-07-06. Failing to See or Failing to Know? Attributing Errors in Vision-Language Models. https://arxiv.org/abs/2607.04683
Cite the original work for its findings. Save a collection to share your selection of sources.