SearcharxivSearch

arXiv subjects

Arman Rahmim

Publications and source records attributed to Arman Rahmim.

At least 19 recordsLinked to original sources

Volumetric Radiology AI in the Era of Multimodal Large Language Models

Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual representations, or report-derived text. Reliable volumetric radiology AI therefore requires representations that preserve task-relevant three-dimensional (3D) information and systems that can access, verify, and integrate this information across clinical workflows. In this Review, we examine more than 200 publications through July 2026. We organize the literature around volumetric representation and multimodal understanding at the model level, agentic orchestration at the system level, and their links to clinical applications and evaluation. We review volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction. We distinguish settings in which selected 2D views or report-mediated reasoning may suffice from those that warrant native volumetric modeling. We also introduce a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation. Across the literature, native volumetric modeling and agentic capabilities depend on the spatial, quantitative, contextual, and workflow requirements of the intended task. Clinical credibility requires faithful volumetric representation, traceable system behavior, claim-aligned validation, and clearly defined human oversight in realistic workflows.

cs.AI

Rethinking Artificial Intelligence in Medical Imaging: Assumptions, Reality, and Reframing

Medical imaging has served as primary proving ground for clinical artificial intelligence (AI), yet a decade of intense research has not translated into proportionate bedside impact. We argue that this gap is not primarily a product of insufficient algorithmic performance, inadequate regulation, or limited explainability. Rather, it reflects a structural misalignment, between how AI systems are designed and evaluated, and how clinical decisions are made. This Perspective identifies six interconnected dimensions of this misalignment: the dominance of pixel-only models in a multimodal clinical world; the erosion of physician trust through opaque and inflexible systems; the unfulfilled promise of foundation models in data-sparse medical domains; the persistent bottleneck of non-shareable, under-curated datasets; the gap between validated algorithms and deployable clinical platforms; and the failure of prediction-centric AI to generate actionable clinical guidance. For each dimension, we reframe the problem and propose a path forward, culminating in a vision of agentic, physician-aligned AI that extends, rather than replaces, clinical judgment.

physics.med-ph

HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT

Head and neck cancers (HNC) represent a significant global health burden, with accurate tumor delineation being essential for effective radiotherapy planning. The complexity of the oropharyngeal anatomy, combined with the heterogeneous appearance of tumors on imaging, makes manual segmentation time-intensive and subject to inter-observer variability. Beyond segmentation, predicting long-term clinical outcomes, such as recurrence-free survival (RFS), and determining human papillomavirus (HPV) status from noninvasive imaging, remain challenging yet clinically valuable goals. The HECKTOR 2025 challenge addresses these needs by establishing a comprehensive benchmark for automated HNC analysis using multimodal PET/CT imaging and electronic health records. Building on previous editions (2020-2022), this challenge features an expanded multi-institutional dataset comprising over 1,100 patients from 10 centers worldwide. Participants were tasked with three complementary objectives: (1) segmenting primary gross tumor volumes (GTVp) and metastatic lymph nodes (GTVn), (2) predicting recurrence-free survival, and (3) classifying HPV status. The challenge attracted 35 registered teams, with 15 final submissions evaluated on a held-out test set. Top-performing algorithms achieved a mean Dice similarity coefficient of 0.75 for segmentation, a concordance index of 0.66 for survival prediction, and a balanced accuracy of 0.56 for HPV classification. This paper presents a comprehensive analysis of the submitted methodologies, evaluates their performance across different lesion characteristics, and discusses their implications for clinical translation in automated oncology workflows and decision support systems.

cs.CV

Radiuma: A Unified Zero-Code Executable Graphical Workflow Generator for Reproducible and Shareable Medical Image Analysis and Machine Learning

Medical image computing software is essential for identifying imaging biomarkers that can support diagnosis, prognosis, treatment planning, and clinical research. However, the lack of standardized, user-friendly, and reproducible software environments has limited the broader adoption of advanced medical image analysis workflows. We present Radiuma, a freely available modular platform designed to support reliable and reproducible medical image analysis across multiple modalities and file formats. Radiuma integrates image reading, visualization, registration, fusion, processing, segmentation, radiomics feature extraction, and machine learning modules for classification, regression, and clustering. Its modular design allows users to execute each component independently or connect modules through a visual workflow system, where the output of one step can be graphically passed to the next. This enables the creation of custom, executable, and reproducible multi-step pipelines without requiring extensive programming expertise. Results from each module can be inspected directly in the visualization window, providing immediate feedback on processing quality and workflow accuracy. Radiuma also supports saving and sharing customized workflows, promoting transparency, reusability, and consistency across collaborative studies. By combining flexibility, usability, and standardized analysis tools, Radiuma provides a practical environment for radiomics and machine learning research in clinical and translational settings. The platform is designed to be accessible to users with diverse expertise, including radiologists, physicists, clinicians, and data scientists.

cs.CV

Renal Blood Flow Quantification During Standard Myocardial Perfusion Imaging with Rubidium-82 Positron Emission Tomography

Background: Renal blood flow (RBF) is an important marker of kidney health, but noninvasive assessment is not routinely used in clinical imaging. We evaluated the feasibility and physiologic validity of quantifying renal transport of Rubidium-82 (K1) during standard myocardial perfusion imaging (MPI) PET. Methods: We studied 126 patients (age 60 +/- 12 years; 48% male; 51% Black) undergoing clinically indicated rest and stress Rb-82 MPI, in whom at least one kidney was partially visualized within the axial field of view. Volumes of interest were drawn over the visible renal cortex. K1 was estimated using a one-tissue compartment model with arterial input functions (AIF) derived from either the left ventricle (LV) or abdominal aorta. Results: LV-derived AIF produced physiologic and internally consistent flow estimates, whereas aorta-derived AIF systematically overestimated K1 and flow. LV-based measurements were therefore used for all analyses. K1 demonstrated nonlinear flow dependence consistent with the Renkin-Crone extraction model, plateauing at higher perfusion states. Renal K1 and flow declined progressively with worsening kidney function, from 1.24 +/- 0.35 ml/min/g (eGFR >= 60) to 0.53 +/- 0.22 ml/min/g (eGFR < 15; P < 0.0001). Only patients with preserved eGFR showed significant hyperemic augmentation. ROC analysis demonstrated excellent discrimination for reduced kidney function (AUC > 0.90). Conclusion: Opportunistic renal K1 quantification during routine Rb-82 PET is feasible, physiologically consistent, and strongly associated with kidney function.

physics.med-ph

Robust Multicenter CT Radiogenomics for Dual EGFR and KRAS Prediction in Lung Cancer with Stability-Aware Modeling and SHAP Interpretation

Accurate identification of EGFR and KRAS mutations is essential for precision therapy in non-small cell lung cancer (NSCLC), but tissue genotyping is invasive and may not capture tumor heterogeneity. CT-based radiogenomics offers a noninvasive alternative, although generalization across centers remains challenging. We benchmarked handcrafted radiomics features (HRF), deep feature representations (DFR), and their fusion for three-class mutation prediction (wild-type, KRAS-mutant, and EGFR-mutant) with external testing. We curated 1,023 thoracic CT scans from 12 public datasets across more than 20 centers, including 136 patients with EGFR/KRAS labels. IBSI-compliant HRFs were extracted with standardized preprocessing, and DFRs were derived using PySERA. HRF-only, DFR-only, and fused HRF+DFR pipelines were evaluated using five-fold cross-validation and external testing. A semi-supervised pseudo-labeling strategy leveraged unlabeled CT scans, and SHAP supported interpretability. In external testing, HRF-based models generalized best, achieving AUC 0.77 +/- 0.07 and accuracy 0.77 +/- 0.00. DFR-based models showed a larger drop from cross-validation to external testing, with best external AUC around 0.57 +/- 0.05. Fusion improved robustness over DFR-only models but did not consistently outperform HRFs. SHAP identified morphology- and heterogeneity-related radiomic phenotypes as key predictors. Standardized handcrafted radiomics within a multicenter semi-supervised framework may provide a generalizable and interpretable approach for CT-based EGFR/KRAS stratification.

physics.med-ph

Towards Routine AI-Based PET/CT and SPECT/CT Lesion Segmentation and Tracking in PSMA Theranostics

Quantitative molecular imaging is central to treatment response assessment in oncology, yet clinical practice remains largely dominated by patient-level or limited target-lesion criteria that ignore inter-lesion heterogeneity. This limitation is particularly important in prostate cancer, where PSMA PET/CT can reveal extensive skeletal and nodal metastatic disease that often evolves heterogeneously under therapy. Accurate and scalable lesion segmentation and tracking across serial PSMA PET/CT and post-therapy SPECT/CT scans is therefore essential for implementing emerging PSMA-specific response frameworks, such as RECIP 1.0, and for enabling lesion-level dosimetry in 177Lu-PSMA radiopharmaceutical therapies (RPTs). This article examines clinical motivations, technical foundations, and future pathways for automated lesion tracking in prostate cancer imaging. We focus on the unique requirements introduced by PSMA PET/CT compared with FDG PET/CT and highlight the critical role of quantitative SPECT/CT in linking imaging-derived disease characterization with delivered therapeutic dose. Recent advances in AI-based segmentation and automated lesion matching now make scalable longitudinal lesion correspondence feasible, providing comprehensive infrastructure for standardized response assessment and personalized PSMA-based theranostics

physics.med-ph

A Clinically Anchored Radiomics Dictionary for Explainable TI-RADS-Based Thyroid Nodule Classification in Ultrasound; Dictionary Version TU1.0

Artificial intelligence based radiomics models for thyroid ultrasound (US) often achieve strong diagnostic performance but remain difficult to interpret, limiting clinical trust and adoption. We developed and validated an interpretable radiomic feature (RF) framework for thyroid nodule classification by linking quantitative US features to the Thyroid Imaging Reporting and Data System (TI-RADS) semantic lexicon through a clinically grounded radiomics dictionary. The dictionary mapped TI-RADS categories, including composition, echogenicity, shape, margin, and echogenic foci, to Image Biomarker Standardization Initiative compliant RFs extracted from two-dimensional US images. Relationships were defined through expert consensus and examined using Shapley Additive Explanations (SHAP). Three multicenter datasets were combined, yielding 5,542 nodules. A total of 107 RFs were extracted using PyRadiomics and normalized with min-max scaling. For benign versus malignant classification, 27 feature selection methods were paired with 25 classifiers and evaluated using stratified five-fold cross-validation on 70% of the data, followed by testing on the remaining 30%. Robust model selection used a stability-aware composite score combining mean performance and variability across balanced accuracy, precision, recall, F1-score, and ROC-AUC. The proposed dictionary enabled direct interpretation of radiomic signatures in TI-RADS terms. The best model, Select-From-Model based on logistic regression with Extra-Trees, achieved a test ROC-AUC of 0.941 +/- 0.005. SHAP analysis showed that texture heterogeneity was the dominant malignancy signal, with gray level run length matrix non-uniformity, intensity dispersion, and kurtosis aligning with high-risk TI-RADS descriptors. These findings support transparent and clinically meaningful thyroid nodule risk stratification from US.

physics.med-ph

Beyond the Tumor: Recurrence-Prone Radiomics for Prognostication in Negative PSMA PET/CT scans of Prostate Cancer

In patients with biochemical recurrence of prostate cancer and negative PSMA PET/CT, radiomics features extracted from recurrence-prone organs can predict clinical progression and progression-free survival. In a cohort of 132 patients, combining PET/CT radiomics with clinical variables significantly improved prognostic performance (C-index 0.74 vs. 0.65). Model performance was influenced by reader diagnostic certainty and remained robust on external validation. These findings suggest that radiomics from visually negative scans capture subclinical disease and provide added value for risk stratification.

physics.med-ph

Multi-Kernel Gated Decoder Adapters for Robust Multi-Task Thyroid Ultrasound under Cross-Center Shift

Thyroid ultrasound (US) automation couples two competing requirements: global, geometry-driven reasoning for nodule delineation and local, texture-driven reasoning for malignancy risk assessment. Under cross-center domain shift, these cues degrade asymmetrically, yet most multi-task pipelines rely on a single shared backbone, often inducing negative transfer. In this paper, we characterize this interference across CNN (ResNet34) and medical ViT (MedSAM) backbones, and observe a consistent trend: ViTs transfer geometric priors that benefit segmentation, whereas CNNs more reliably preserve texture cues for malignancy discrimination under strong shift and artifacts. Motivated by this failure mode, we propose a lightweight family of decoder-side adapters, the Multi-Kernel Gated Adapter (MKGA) and a residual variant (ResMKGA), which refine multi-scale skip features using complementary receptive fields and apply semantic, context-conditioned gating to suppress artifact-prone content before fusion. Across two US benchmarks, the proposed adapters improve cross-center robustness: they strengthen out-of-domain segmentation and, in the CNN setting, yield clear gains in clinical TI-RADS diagnostic accuracy compared to standard multi-task baselines. Code and models will be released.

cs.CV

Virtual Theranostic Trials: New Approach Methodologies (NAMs) and Precision Radiopharmaceutical Therapies

Radiopharmaceutical therapies are expanding rapidly, but clinical evidence generation is limited by operational constraints and biological heterogeneity in radiopharmaceutical delivery and radiation risk. Virtual theranostic trials can run as companion trials, linking quantitative imaging to patient-specific models to generate evidence for personalized injections and scheduling, accelerate development, and broaden access.

physics.med-ph

Unveiling and Bridging the Functional Perception Gap in MLLMs: Atomic Visual Alignment and Hierarchical Evaluation via PET-Bench

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in tasks such as abnormality detection and report generation for anatomical modalities, their capability in functional imaging remains largely unexplored. In this work, we identify and quantify a fundamental functional perception gap: the inability of current vision encoders to decode functional tracer biodistribution independent of morphological priors. Identifying Positron Emission Tomography (PET) as the quintessential modality to investigate this disconnect, we introduce PET-Bench, the first large-scale functional imaging benchmark comprising 52,308 hierarchical QA pairs from 9,732 multi-site, multi-tracer PET studies. Extensive evaluation of 19 state-of-the-art MLLMs reveals a critical safety hazard termed the Chain-of-Thought (CoT) hallucination trap. We observe that standard CoT prompting, widely considered to enhance reasoning, paradoxically decouples linguistic generation from visual evidence in PET, producing clinically fluent but factually ungrounded diagnoses. To resolve this, we propose Atomic Visual Alignment (AVA), a simple fine-tuning strategy that enforces the mastery of low-level functional perception prior to high-level diagnostic reasoning. Our results demonstrate that AVA effectively bridges the perception gap, transforming CoT from a source of hallucination into a robust inference tool and improving diagnostic accuracy by up to 14.83%. Code and data are available at https://github.com/yezanting/PET-Bench.

cs.CV

Towards Interpretable AI in Personalized Medicine: A Radiological-Biological Radiomics Dictionary Connecting Semantic Lung-RADS and imaging Radiomics Features; Dictionary LC 1.0

Lung cancer remains the leading cause of cancer-related mortality worldwide, with survival strongly dependent on early detection. Standard-dose computed tomography (CT) screening using the Lung Imaging Reporting and Data System (Lung-RADS) standardizes pulmonary nodule assessment but is limited by inter-reader variability and reliance on qualitative descriptors, while radiomics offers quantitative biomarkers that often lack clinical interpretability. To bridge this gap, we propose a radiological-biological dictionary that aligns radiomic features (RFs) with Lung-RADS semantic categories. A clinically informed dictionary translating ten Lung-RADS descriptors into radiomic proxies was developed through literature curation and validated by eight expert reviewers. As a proof of concept, imaging and clinical data from 977 patients across 12 collections in The Cancer Imaging Archive (TCIA) were analyzed; following preprocessing and manual segmentation, 110 RFs per nodule were extracted using PyRadiomics in compliance with the Image Biomarker Standardization Initiative (IBSI). A semi-supervised learning framework incorporating 499 labeled and 478 unlabeled cases was applied to improve generalizability, evaluating seven feature selection methods and ten interpretable classifiers. The optimal pipeline (ANOVA feature selection with a support vector machine) achieved a mean validation accuracy of 0.79. SHapley Additive exPlanations (SHAP) analysis identified key RFs corresponding to Lung-RADS semantics such as attenuation, margin irregularity, and spiculation, supporting the validity of the proposed mapping. Overall, this dictionary provides an interpretable framework linking radiomics and Lung-RADS semantics, advancing explainable artificial intelligence for CT-based lung cancer screening.

physics.med-ph

Towards Integrated Clinical-Computational Nuclear Medicine

The field of Clinical-Computational Nuclear Medicine is rapidly advancing, fueled by AI, tracer kinetic modeling, radiomics, and integrated informatics. These technologies improve imaging quality, automate lesion detection, and enable personalized radiopharmaceutical therapy through physiologically based pharmacokinetic (PBPK) modeling and voxel-level dosimetry. Workflow automation and Natural Language Processing (NLP) further enhance operational efficiency. However, successful implementation and adoption of these tools require clinical oversight to ensure accuracy, interpretability, and patient safety. This paper highlights key computational innovations and emphasizes the critical role of clinician-guided evaluation in shaping the future of precision imaging and therapy.

physics.med-ph

PySERA: Open-Source Standardized Python Library for Automated, Scalable, and Reproducible Handcrafted and Deep Radiomics

Radiomics enables the extraction of quantitative biomarkers from medical images for precision modeling, but reproducibility and scalability remain limited due to heterogeneous software implementations and incomplete adherence to standards. Existing tools also lack unified support for deep learning based radiomics. To address these limitations, we introduce PySERA, an open source, Python native, standardized radiomics framework designed for automation, reproducibility, and seamless AI integration. PySERA reimplements the MATLAB based SERA platform within a modular, object oriented architecture and computes 557 features, including 487 Image Biomarker Standardization Initiative (IBSI) compliant features, 10 moment invariant descriptors, and 60 diagnostic features, together with deep learning radiomics embeddings from pretrained networks such as ResNet50, DenseNet121, and VGG16. The framework provides standardized preprocessing, including resampling, discretization, and normalization, multi format image input and output, adaptive memory management, and a parallel multicores extraction engine. PySERA integrates natively with major machine learning ecosystems including scikit learn, PyTorch, TensorFlow, MONAI and others. In IBSI benchmarks, PySERA achieved more than 94 percent reproducibility and outperformed PyRadiomics while closely matching MITK. Across eight public datasets, it achieved predictive accuracies ranging from 0.54 to 0.87, consistently exceeding PyRadiomics. PySERA unifies handcrafted and deep learning radiomics in a transparent, scalable, and extensible Python framework, establishing a robust foundation for reproducible, AI ready imaging research.

physics.med-ph

Mathematical and Computational Nuclear Oncology: Toward Optimized Radiopharmaceutical Therapy via Digital Twins

This article presents the general framework of theranostic digital twins (TDTs) in computational nuclear medicine, designed to support clinical decision-making and improve cancer patient prognosis through personalized radiopharmaceutical therapies (RPTs). It outlines potential clinical applications of TDTs and proposes a roadmap for successful implementation. Additionally, the chapter provides a conceptual overview of the current state of the art in the mathematical and computational modeling of RPTs, highlighting key challenges and the strategies being pursued to address them.

q-bio.OT

Predictive Dosimetry in PSMA-Targeted Radiopharmaceutical Therapies: A PBPK Modeling and Machine Learning Study

Predictive dosimetry is central to enabling personalized radiopharmaceutical therapy (RPT), particularly in prostate specific membrane antigen (PSMA) targeted theranostics. In this work, we develop a three layer computational framework that integrates physiologically based pharmacokinetic (PBPK) modeling with machine learning (ML) to predict both physical (AUC, absorbed dose) and biological (BED, EQD2) dosimetric endpoints in tumors and major organs. In the first layer, we generated 640 virtual patients using PBPK simulations of F-18, Ga-68, and Cu-64 labeled PSMA PET tracers paired with Lu-177 PSMA therapy, producing 15360 tumor and organ time activity curves (TACs) under realistic biological variability and PET-like noise. In the second layer, TACs were transformed into quantitative kinetic features and mapped to physical and biological dose metrics. In the third layer, ML models (Random Forest, Extra Trees, Ridge, Gradient Boosting, and XGBoost) were trained to predict RPT doses from PET derived features, with performance evaluated using mean absolute percentage error (MAPE) and R2. Cu-64 PSMA-617 based PET yielded the most robust predictions, achieving tumor dose MAPE as low as 8 percent and 10 to 20 percent for normal organs, while F-18 DCFPyL showed volume dependent performance and Ga-68 PSMA-11 exhibited higher variability. SHAP analysis revealed that peak uptake, clearance, and early kinetic features dominated predictive performance across organs and endpoints. This PBPK ML framework enables scalable, physiology informed predictive dosimetry and provides a foundation for trial design and patient specific treatment planning in PSMA targeted RPT. These results demonstrate that pre therapy PET can serve as a reliable surrogate for post therapy dosimetry, enabling scalable personalization of PSMA targeted RPT.

physics.med-ph

Click, Predict, Trust: Clinician-in-the-Loop AI Segmentation for Lung Cancer CT-Based Prognosis within the Knowledge-to-Action Framework

Lung cancer remains the leading cause of cancer mortality, with CT imaging central to screening, prognosis, and treatment. Manual segmentation is variable and time-intensive, while deep learning (DL) offers automation but faces barriers to clinical adoption. Guided by the Knowledge-to-Action framework, this study develops a clinician-in-the-loop DL pipeline to enhance reproducibility, prognostic accuracy, and clinical trust. Multi-center CT data from 999 patients across 12 public datasets were analyzed using five DL models (3D Attention U-Net, ResUNet, VNet, ReconNet, SAM-Med3D), benchmarked against expert contours on whole and click-point cropped images. Segmentation reproducibility was assessed using 497 PySERA-extracted radiomic features via Spearman correlation, ICC, Wilcoxon tests, and MANOVA, while prognostic modeling compared supervised (SL) and semi-supervised learning (SSL) across 38 dimensionality reduction strategies and 24 classifiers. Six physicians qualitatively evaluated masks across seven domains, including clinical meaningfulness, boundary quality, prognostic value, trust, and workflow integration. VNet achieved the best performance (Dice = 0.83, IoU = 0.71), radiomic stability (mean correlation = 0.76, ICC = 0.65), and predictive accuracy under SSL (accuracy = 0.88, F1 = 0.83). SSL consistently outperformed SL across models. Radiologists favored VNet for peritumoral representation and smoother boundaries, preferring AI-generated initial masks for refinement rather than replacement. These results demonstrate that integrating VNet with SSL yields accurate, reproducible, and clinically trusted CT-based lung cancer prognosis, highlighting a feasible path toward physician-centered AI translation.

cs.CV