arXiv · 2609.26093
RECAP: Relation Evidence Calibration for Detecting Spatial Relation Hallucinations in Vision-Language Models
Abstract
Vision-language models can answer spatial relation questions confidently even when the image supports an incompatible relation. We formulate relation-grounded selective prediction: accept or reject an already-produced yes/no answer by auditing its visual support, rather than treating uncertainty as evidence. RECAP, our relation-evidence calibration framework, compares image-conditioned likelihoods for a claim, its semantic contradictions, and optional one-sided supports, then converts these witnesses into an answer-conditioned rejection risk. A calibration-only gate preserves confidence as a veto when confidence is demonstrably informative and otherwise deploys relation evidence alone. Across 20 group/image-disjoint splits, RECAP lowers H-FPR@80 over confidence by between 2.0 and 17.9 points on VSR and raises Acc@80 by 3.0, 8.6, and 12.6 points on What'sUp for Qwen3-VL-8B, InternVL3.5-8B, and LLaVA-1.5-7B. It outperforms matched VCD-style visual contrast on all four primary metrics in all six settings. Full-pool VSR fallback, target-ranked GSR-Bench transfer, equal-budget supervised controls, and two additional checkpoints show a consistent operating principle: structured counterevidence complements certainty when confidence is misaligned, while the gate retains confidence when it is already useful.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Feixiang Liu, Qiang Qiu, Qingyang Li, Hui Xu. 2026-08-04. RECAP: Relation Evidence Calibration for Detecting Spatial Relation Hallucinations in Vision-Language Models. https://arxiv.org/abs/2609.26093
Cite the original work for its findings. Save a collection to share your selection of sources.