arXiv · 2606.16667
Look Again Before You Abstain:Budgeted Conformal Evidence Acquisition for Reliable Vision-Language Model
Abstract
Large vision-language models (LVLMs) hallucinate: they assert visual details that the image does not support. A principled remedy is selective prediction with a distribution-free guarantee-verify each claim and abstain when the claim is not grounded, so that the hallucination rate among asserted claims is provably bounded. We show, however, that this guarantee is bought at a brutal price: to keep the hallucination rate below $5\%$ on a balanced object-existence benchmark, a state-of-the-art conformal filter must abstain on more than $80\%$ of claims. We argue that abstention is wasteful when more visual evidence is cheaply available, and introduce Budgeted Conformal Evidence Acquisition (BCEA), which replaces the binary answer/abstain decision with a three-way choice: answer, abstain, or acquire additional visual evidence by re-examining the image (zooming, cropping, or applying a claim-specific intervention) under a bounded
Explore related subjects
Keep this discovery
Jian Xu, Yanning Wu, Delu Zeng, John Paisley, Qibin Zhao. 2026-06-15. Look Again Before You Abstain:Budgeted Conformal Evidence Acquisition for Reliable Vision-Language Model. https://arxiv.org/abs/2606.16667
Cite the original work for its findings. Save a collection to share your selection of sources.