arXiv · 2609.29358
Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models
Abstract
Vision-language models such as CLIP achieve strong zero-shot classification, yet under distribution shift, visual embeddings drift from fixed text embeddings. Training-free calibration avoids the per-sample optimization of prompt learning, but prior feature calibration gives each image the full bias of one hard cluster. We propose Domain Recentering with Confidence Calibration (DRC), a training-free method adapting CLIP from a set of unlabeled target images. DRC fits a Gaussian mixture once and subtracts from each embedding a posterior-weighted average of component means. It then removes residual class preference with a log-prior correction, estimating the prior from confidence-weighted predictions. Among compared methods, DRC achieves the highest average accuracy on cross-domain datasets, exceeding zero-shot CLIP by 4.13 and 5.07 points with ViT-B/16 and ResNet-50, with gains over CLIP also holding under ImageNet distribution shifts.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Youngeun Seol, Jimin Shin, Heeseo Yoon, Uiwon Hwang. 2026-09-24. Domain Recentering and Confidence-Weighted Prior Calibration for Vision-Language Models. https://arxiv.org/abs/2609.29358
Cite the original work for its findings. Save a collection to share your selection of sources.