arXiv · 2505.20122
MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models
Abstract
This paper introduces MEBench, a novel benchmark for evaluating mutual exclusivity (ME) bias, a cognitive phenomenon observed in children during word learning. Unlike traditional ME tasks, MEBench further incorporates spatial reasoning to create more challenging and realistic evaluation settings. To facilitate controlled experimentation, we also present a flexible and scalable data generation pipeline that supports the construction of diverse annotated scenes. We assess the performance of various vision-language models (VLMs) on this benchmark using novel evaluation metrics that capture key aspects of ME-based reasoning. We find that these VLMs exhibit weak ME bias, while showing some ability to leverage extra spatial context to resolve ambiguity in multiple novel object settings. Project page: http://mebench.github.io/.
Explore related subjects
Keep this discovery
Anh Thai, Stefan Stojanov, Zixuan Huang, Bikram Boote, James M. Rehg. 2025-05-26. MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models. https://arxiv.org/abs/2505.20122
Cite the original work for its findings. Save a collection to share your selection of sources.