arXiv · 2511.03908
Context informs pragmatic interpretation in vision-language models
Abstract
Iterated reference games - in which players repeatedly pick out novel referents using language - present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision-language models on trials from iterated reference games, varying the given context in terms of amount, order, and relevance. Without relevant context, models were above chance but substantially worse than humans. However, with relevant context, model performance increased dramatically over trials. Few-shot reference games with abstract referents remain a difficult task for machine learning models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce, Michael C. Frank. 2025-11-05. Context informs pragmatic interpretation in vision-language models. https://arxiv.org/abs/2511.03908
Cite the original work for its findings. Save a collection to share your selection of sources.