arXiv · 2609.22008
DiaVLo: Diagnosing Behaviours of Vision-Language Models
Abstract
Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs' generation capabilities to construct specifications of desired and observed VLM behaviours, surfacing potential misalignments. Beyond this, DiaVLo also provides causal estimates to identify the most influential concepts steering VLM behaviours. We evaluate DiaVLo on several open-source VLMs under both classification and generation conditions. Our experiments show that DiaVLo produces behaviour labels that correlate with model performance and provide context for measured performance. DiaVLo surfaced behaviours that are clearly aligned and misaligned, alongside patterns in how VLMs perceive, organise, and prioritise concepts.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lorenzo Corti, Jie Yang. 2026-09-18. DiaVLo: Diagnosing Behaviours of Vision-Language Models. https://arxiv.org/abs/2609.22008
Cite the original work for its findings. Save a collection to share your selection of sources.