arXiv · 2609.23860
Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models
Abstract
Multimodal Large Language Models (MLLMs) commonly reuse visual encoders pretrained with CLIP, although the features of these ViTs are ultimately consumed by autoregressive LLMs. We refer to this mismatch as the semantic-interface gap and introduce MedMLIP, a framework that pretrains the visual encoder through report generation with a frozen LLM, while employing Local Relational Distillation (LRD) to preserve relationships among visual patches to avoid visual collapse. We pretrain MedMLIP on IU-Xray and Open-PMC-300K and evaluate the resulting encoders on VQA-RAD and SLAKE. Only the ViT is transferred, while the guiding LLM and projector are replaced, allowing us to assess cross-LLM transferability. Our cross-LLM transfer experiments demonstrate the value of pretraining visual encoders for their autoregressive LLM interface while trying to preserve more fine-grained visual information. Code and the pretrained model are available at https://github.com/SkyCol/MedMLIP
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tianyou Jiang. 2026-09-20. Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models. https://arxiv.org/abs/2609.23860
Cite the original work for its findings. Save a collection to share your selection of sources.