arXiv · 2609.37230
Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry
Abstract
Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim. 2026-09-29. Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry. https://arxiv.org/abs/2609.37230
Cite the original work for its findings. Save a collection to share your selection of sources.