arXiv · 2601.15891
RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture
Abstract
Vision-language pretraining has driven progress in medical image representation learning, but it depends on paired image-text data and can inherit reporting bias from clinical narratives. We study whether language-free predictive pretraining can produce an image encoder that transfers effectively to radiology report generation. RadJEPA is a chest-X-ray adaptation of I-JEPA, pretrained on approximately 840K unlabeled radiographs using latent context-to-target prediction. Our primary contribution is an extensive empirical evaluation of this language-free encoder for report generation: the frozen image encoder is coupled to a trainable two-layer projector and language decoder, and is also substituted into four established vision-language backbones. Across MIMIC-CXR and IU-Xray, RadJEPA matches or exceeds the evaluated image-only and image-text baselines on lexical, entity-relation, and clinical-label metrics. Controlled MIMIC-only comparisons provide evidence that the predictive objective contributes beyond domain-specific pretraining, while broader comparisons also reflect differences in pretraining data, model capacity, and input resolution. Complementary classification and segmentation experiments assess transfer beyond report generation.
Explore related subjects
Keep this discovery
Anas Anwarul Haq Khan, Mariam Husain, Pratik Jalan, Kshitij Jadhav. 2026-01-22. RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture. https://arxiv.org/abs/2601.15891
Cite the original work for its findings. Save a collection to share your selection of sources.