SearcharxivSearch

arXiv · 2504.10916

Embedding Radiomics into Vision Transformers for Multimodal Medical Image Classification

Abstract

Background: Deep learning has significantly advanced medical image analysis, with Vision Transformers (ViTs) offering a powerful alternative to convolutional models by modeling long-range dependencies through self-attention. However, ViTs are inherently data-intensive and lack domain-specific inductive biases, limiting their applicability in medical imaging. In contrast, radiomics provides interpretable, handcrafted descriptors of tissue heterogeneity but suffers from limited scalability and integration into end-to-end learning frameworks. In this work, we propose the Radiomics-Embedded Vision Transformer (RE-ViT) that combines radiomic features with data-driven visual embeddings within a ViT backbone. Purpose: To develop a hybrid RE-ViT framework that integrates radiomics and patch-wise ViT embeddings through early fusion, enhancing robustness and performance in medical image classification. Methods: Following the standard ViT pipeline, images were divided into patches. For each patch, handcrafted radiomic features were extracted and fused with linearly projected pixel embeddings. The fused representations were normalized, positionally encoded, and passed to the ViT encoder. A learnable [CLS] token aggregated patch-level information for classification. We evaluated RE-ViT on three public datasets (including BUSI, ChestXray2017, and Retinal OCT) using accuracy, macro AUC, sensitivity, and specificity. RE-ViT was benchmarked against CNN-based (VGG-16, ResNet) and hybrid (TransMed) models. Results: RE-ViT achieved state-of-the-art results: on BUSI, AUC=0.950+/-0.011; on ChestXray2017, AUC=0.989+/-0.004; on Retinal OCT, AUC=0.986+/-0.001, which outperforms other comparison models. Conclusions: The RE-ViT framework effectively integrates radiomics with ViT architectures, demonstrating improved performance and generalizability across multimodal medical image classification tasks.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhenyu Yang, Haiming Zhu, Rihui Zhang, Haipeng Zhang, Jianliang Wang, Chunhao Wang, Minbin Chen, Fang-Fang Yin. 2025-04-15. Embedding Radiomics into Vision Transformers for Multimodal Medical Image Classification. https://arxiv.org/abs/2504.10916

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Phase-contrast micro-CT for intra-operative breast tumour margin assessment using a microfocus x-ray source and photon-counting detector

Objective: Intra-operative tumour margin assessment during breast-conserving surgery requires rapid, high-resolution imaging of excised tissue, allowing the surgical team to take appropriate action within a single operation. This study evaluates a custom propagation-based phase-contrast micro-computed tomography (micro-CT) system designed to meet these clinical constraints without specialised optical elements. Methods: The experimental setup pairs a microfocus x-ray source with a photon-counting detector in a cone-beam geometry. We explore how the spatial coherence of the source can provide propagation-based phase contrast -- with no additional specialised optical elements -- and balance this against maximising the x-ray flux of the cone-beam geometry. System performance was evaluated across two anode target materials and filtration configurations at various tube power settings. Imaging capabilities were validated using anthropomorphic breast tissue phantoms and a formalin-fixed paraffin-embedded (FFPE) breast tissue specimen, with reconstructions compared against gold-standard histology. Results: An unfiltered tungsten target operated at 40 kVp yielded optimal image quality. The optimised system achieved high-resolution CT reconstructions of a 5 cm diameter sample with an isotropic voxel size of 40.7 $\upmu\text{m}$ in a scan time of 12 minutes. Reconstructed volumes demonstrated strong visual correlation with corresponding histology slides. Conclusion: Combining a microfocus source with a photon-counting detector enables high-resolution, phase-contrast micro-CT within a clinically viable timeframe, demonstrating strong potential for intra-operative margin assessment.

physics.med-ph

Understanding Search and Decision Errors in Liver Metastasis Detection and the Effects of Lower Radiation Dose

The detection performance of liver metastases decreases with the reduction of radiation dose, but misses are heterogeneous. Previous eye tracking work has characterized missed metastases into two categories: search errors i.e., the eyes never land on the lesion, and decision errors i.e., the lesion is seen but not recognized as malignant. We integrated three prior reader studies to answer this question. In all studies, radiologists interpreted the same set of 40 contrast enhanced abdominal CT exams containing 91 liver metastases whose locations had been previously marked. In two studies, the workstation recorded their gaze and eye movements. Using eye dwell times, metastases were classified as search-error-dominant (majority of misses had <2 sec gaze time) or decision-error-dominant (>2 sec gaze time). In the third study, exams were interpreted both at 120 and 200 quality reference mAs (QRM) by ten radiologists. The third study did not include eye tracking. Out of 91 liver metastases, we excluded 16 that were never missed in the eye tracking studies and used 75 liver metastases for the present study.

physics.med-ph

Develop and Optimize 5DCT Imaging Simulation and Reconstruction Methods

Purpose: To develop and optimize a 5DCT (3D + cardiac phase + respiratory phase) imaging simulation and reconstruction pipeline, and to compare two sinogram-space interpolation methods for reconstructing images at arbitrary combinations of cardiac and respiratory phase. Methods: Helical CT projections were simulated from the 4D XCAT phantom across a range of cardiac and respiratory motion states, with Poisson and electronic noise added. Ground-truth-matched volumes were generated at 5 cardiac phases and 10 respiratory amplitudes (50 total phase combinations). Because acquired projections are sparsely and unevenly distributed across this joint phase space, each target slice was reconstructed by interpolating rebinned sinogram rows to the target cardiac phase and respiratory amplitude, using either 2D scattered barycentric interpolation or 2D scattered local linear interpolation with a circular kernel for cardiac phase. Reconstructed volumes were compared to phantom ground truth using mean absolute error (MAE), and to conventional respiratory-gated 4DCT (r4DCT) reconstructed from the same simulated data. Results: Both interpolation methods eliminated the severe axial misalignment artifacts present when helical projections were reconstructed without phase-space interpolation. Local linear interpolation achieved lower MAE than barycentric interpolation across most tested conditions, with the largest improvement at low pitch. The 5DCT pipeline also produced respiratory-only volumes with fewer residual cardiac-motion artifacts than conventional r4DCT reconstructed from the same projection data, including at standard clinical pitch (0.1). Conclusions: 5DCT reconstruction using sinogram-space interpolation is feasible and can jointly resolve cardiac and respiratory motion with better accuracy than conventional 4DCT reconstruction.

physics.med-ph