arXiv · 2603.25383
CLIP-RD: Relational Distillation for Efficient CLIP Knowledge Distillation
Abstract
Contrastive Language-Image Pre-training (CLIP) demonstrates strong zero-shot generalization, but due to substantial computational and memory costs, distillation into lightweight models is required. Existing relational objectives do not explicitly model multidirectional relationships between teacher and student embeddings, potentially leaving the geometric relationships insufficiently constrained. This may disrupt the modality-gap structure important for zero-shot transfer. To address these limitations, we propose a relational distillation framework, CLIP-RD, which introduces two relational methods, Vertical Relational Distillation (VRD) and Cross Relational Distillation (XRD). VRD aligns teacher-student intra-modal similarity distributions to enforce consistent distillation strength across image and text embeddings. Meanwhile, XRD aligns the teacher-image-student-text and teacher-text-student-image similarity distributions to impose bidirectional cross-modal symmetry. By jointly modeling these multidirectional relational structures, CLIP-RD aligns the student's embedding geometry more faithfully to the teacher's, outperforming CLIP-KD by 1.8%p. This performance improvement is maintained across diverse architectures, teacher scales, retrieval tasks, downstream tasks, and corruption settings, with negligible additional training-time overhead.
Explore related subjects
Keep this discovery
Jeannie Chung, Hanna Jang, Ingyeong Yang, Uiwon Hwang, Jaehyeong Sim. 2026-03-26. CLIP-RD: Relational Distillation for Efficient CLIP Knowledge Distillation. https://arxiv.org/abs/2603.25383
Cite the original work for its findings. Save a collection to share your selection of sources.