arXiv · 2503.18931
Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
Abstract
Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this work, we enhance prevailing VFMs through multimodal training, allowing them to effectively process visual inputs at varying resolutions while producing visual representations that are better aligned with language representations, regardless of their original pre-training objectives. To this end, we introduce M-CPT, a Multimodal Continual Pre-Training framework designed to improve the understanding capability of pre-trained VFMs while preserving their strong visual representation quality. M-CPT introduces a Continual Position Embedding (CPE) for handling flexible visual resolutions, along with a feature alignment objective that improves the consistency between visual and textual representations during multimodal training. Extensive experiments on leading VFMs, including DINOv2, SigLIP, and AIMv2, demonstrate that M-CPT consistently improves multimodal understanding performance while preserving strong performance on standard vision benchmarks such as classification and segmentation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yitong Chen, Lingchen Meng, Wujian Peng, Jun Tao, Chenjie Xu, Zuxuan Wu, Yu-Gang Jiang. 2025-03-24. Enhancing Vision Foundation Models via Multimodal Continual Pre-Training. https://arxiv.org/abs/2503.18931
Cite the original work for its findings. Save a collection to share your selection of sources.