arXiv · 2509.03805
Measuring How (Not Just Whether) VLMs Build Common Ground
Abstract
Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is an interactive process in which people gradually develop shared understanding through ongoing communication. We introduce a four-metric suite (grounding efficiency, content alignment, lexical adaptation, and human-likeness) to systematically evaluate VLM performance in interactive grounding contexts. We deploy the suite on 150 self-play sessions of interactive referential games between three proprietary VLMs and compare them with human dyads. All three models diverge from human patterns on at least three metrics, while GPT4o-mini is the closest overall. We find that (i) task success scores do not indicate successful grounding and (ii) high image-utterance alignment does not necessarily predict task success. Our metric suite and findings offer a framework for future research on VLM grounding.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Saki Imai, Mert İnan, Anthony Sicilia, Malihe Alikhani. 2025-09-04. Measuring How (Not Just Whether) VLMs Build Common Ground. https://arxiv.org/abs/2509.03805
Cite the original work for its findings. Save a collection to share your selection of sources.