arXiv · 2605.11107
Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs
Abstract
Vision-language models (VLMs), such as CLIP and SigLIP 2, are widely used for image classification, yet their vision encoders remain vulnerable to systematic biases that undermine robustness. In particular, correlations between foreground objects and their backgrounds constitute a salient and practically important class of spurious dependencies. In this work, we revisit the well-known property of high linear additivity in VLM embedding spaces and show that it enables a decomposition of scene representations into foreground and background components. Leveraging this insight, we introduce a pre-training approach that exploits this property to construct background-invariant representations using synthetic data. Our method achieves, to our knowledge, the first worst-group accuracy exceeding $90\%$ on Waterbirds under perfect ($100\%$) spurious correlation (i.e., no minority-group examples in the training data). Furthermore, it demonstrates strong sim-to-real transfer and requires no access to real-world debiased data, making it practical for real-world deployment.
Explore related subjects
Keep this discovery
Youssef Zaazou, Mark Thomas. 2026-05-11. Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs. https://arxiv.org/abs/2605.11107
Cite the original work for its findings. Save a collection to share your selection of sources.