arXiv · 2509.15330
CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization
Abstract
Recent advances in pre-training vision-language models (VLMs), e.g., contrastive language-image pre-training (CLIP) methods, have shown great potential in learning out-of-distribution (OOD) representations. Despite showing competitive performance, the prompt-based CLIP methods still suffer from: i) inaccurate text descriptions, which leads to degraded accuracy and robustness, and poses a challenge for zero-shot CLIP methods. ii) limited vision-language embedding alignment, which is one important factor affecting generalization performance. To tackle the above issues, this paper proposes a novel Conditional Domain prompt Learning (CoDoL) method, which utilizes readily-available domain information to form prompts and contributes to improved vision-language embedding alignment, which we identify as one factor underlying the observed OOD generalization gains. To capture both instance-specific and domain-specific information, we further propose a lightweight Domain Meta Network (DMN) to generate input-conditional tokens for images in each domain. Extensive experiments on four OOD benchmarks (PACS, VLCS, OfficeHome, and DigitDG) validate the effectiveness of our proposed CoDoL method in terms of empirically improves vision-language embedding alignment across four DG benchmarks, which we present as a contributing factor (rather than the sole cause) of the observed OOD gains.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Min Zhang, Yuyin Wang, Zhongxiang Dai, Zhikang Chen, Jie Zhou, Miao Liu, Sen Cui. 2025-09-18. CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization. https://arxiv.org/abs/2509.15330
Cite the original work for its findings. Save a collection to share your selection of sources.