arXiv · 2605.15019
From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG
Abstract
Multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities (entire images or scenes), creating a mismatch with fine-grained user queries and making failures unverifiable. We introduce GranuVistaVQA, a multimodal benchmark featuring real-world landmarks with element-level annotations across multiple viewpoints, capturing the partial observation challenge where individual images contain only subsets of entities. We further propose GranuRAG, a multi-granularity framework that treats visual elements as first-class retrieval units through three stages: element-level detection and classification, multi-granularity cross-modal alignment for evidence retrieval, and attribution-constrained generation. By grounding retrieval at the element level rather than relying on implicit attention, our approach enables transparent error diagnosis. Experiments demonstrate that GranuRAG achieves up to 29.2% improvement over six strong baselines for this task.
Explore related subjects
Keep this discovery
Guanhua Chen, Chuyue Huang, Yutong Yao, Shudong Liu, Xueqing Song, Lidia S. Chao, Derek F. Wong. 2026-05-14. From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG. https://arxiv.org/abs/2605.15019
Cite the original work for its findings. Save a collection to share your selection of sources.