arXiv · 2605.03866
Memory-Efficient Continual Learning with CLIP Models
Abstract
Contrastive Language-Image Pretraining (CLIP) models excel at understanding image-text relationships but struggle with adapting to new data without forgetting prior knowledge. To address this, models are typically fine-tuned using both new task data and a memory buffer of past tasks. However, CLIP's contrastive loss suffers when the memory buffer is small, leading to performance degradation on previous tasks. We propose a memory-efficient, distributionally robust method that dynamically reweights losses per class during training. Our approach, tested on class incremental settings (CIFAR-100, ImageNet1K) and a domain incremental setting (DomainNet) adapts CLIP models quickly while minimizing catastrophic forgetting, even with minimal memory usage.
Explore related subjects
Keep this discovery
Ryan King, Gang Li, Bobak Mortazavi, Tianbao Yang. 2026-05-05. Memory-Efficient Continual Learning with CLIP Models. https://arxiv.org/abs/2605.03866
Cite the original work for its findings. Save a collection to share your selection of sources.