SearcharxivSearch

arXiv subjects

Theodoros N. Arvanitis

Publications and source records attributed to Theodoros N. Arvanitis.

12 recordsLinked to original sources

Exploiting Completeness Perception with Diffusion Transformer for Unified 3D MRI Synthesis

Missing data problems, such as missing modalities in multi-modal brain MRI and missing slices in cardiac MRI, pose significant challenges in clinical practice. Existing methods rely on external guidance to supply detailed missing-state information for instructing generative models to synthesize missing MRIs. However, manual indicators are not always available or reliable in real-world scenarios due to the unpredictable nature of clinical environments. Moreover, these explicit masks are not informative enough to provide guidance for improving semantic consistency. In this work, we argue that generative models should infer and recognize missing states in a self-perceptive manner, enabling them to better capture subtle anatomical and pathological variations. Towards this goal, we propose CoPeDiT, a shared completeness-perception framework for 3D MRI synthesis, following a common conditioning strategy with task-specific instantiations for different missing-data scenarios. Specifically, we incorporate dedicated pretext tasks into our tokenizer, CoPeVAE, empowering it to learn completeness-aware discriminative prompt tokens, and design MDiT3D, a specialized diffusion transformer architecture for 3D MRI synthesis that effectively uses the completeness-aware prompt tokens as guidance to enhance semantic consistency in 3D space. Comprehensive evaluations on three large-scale MRI datasets demonstrate that CoPeDiT consistently improves upon state-of-the-art methods across diverse missing patterns, yielding high-fidelity and structurally consistent MRI synthesis. Our code is available at https://github.com/JK-Liu7/CoPeDiT.

eess.IV

Chain of Flow: ECG-Conditioned 4D Cardiac Cine Generation from Patient-Specific Anatomical Anchor

Cardiac cine magnetic resonance imaging (MRI) is central to functional cardiac assessment, yet a full current cine sequence may not always be directly available at the point of analysis. We introduce Chain of Flow (COF), an electrocardiography (ECG)-conditioned framework that combines patient-specific MRI and current ECG for subject-specific 4D cardiac cine generation. On the UK Biobank dataset, COF achieves strong image-level fidelity and downstream function-oriented performance on a shared same-visit evaluable benchmark. Multi-slice and multi-resolution analyses indicate stable structural generation quality across the short-axis stack and heterogeneous acquisition resolutions. Controlled phase-robustness analyses across resampled input MRI phases further provide same-visit proxy support for patient-specific MRI plus current ECG when a target MRI phase is not directly observed. A cross-visit route provides exploratory serial evidence, with the clearest gains in current-facing region-of-interest readout. Disease-category functional audits, case-level volume-trajectory evidence review further delineate where the current patient-specific MRI plus ECG formulation remains stable for anatomy-aware downstream cardiac analysis. Code is available at https://anonymous.4open.science/r/COF-paper-release-C88B.

cs.CV

Hi-GaTA: Hierarchical Gated Temporal Aggregation Adapter for Surgical Video Report Generation

Automated, clinician-grade assessment reports for surgical procedures could reduce documentation burden and provide objective feedback, yet remain challenging due to the difficulty of aligning dense spatio-temporal video representations with language-based reasoning and the scarcity of high-quality, privacy-preserving datasets. To address this gap, we establish a benchmark comprising 214 high-quality simulated surgical videos paired with surgeon-authored evaluation reports. Building on this resource, we propose a Perception-Alignment-Reasoning framework for surgical video report generation, featuring Hi-GaTA, a novel lightweight temporal adapter that efficiently compresses long video sequences into compact, LLM-compatible visual prefix tokens through short-to-long-range temporal aggregation. For robust visual perception, we pretrain Sur40k, a surgical-specific ViViT-style video encoder on 40,000 minutes of public surgical videos to capture fine-grained spatio-temporal procedural priors. Hi-GaTA employs a temporal pyramid with text-conditioned dual cross-attention, and improves multi-scale consistency through cross-level gated fusion and an increasing-depth strategy. Finally, we fine-tune the LLM backbone using LoRA to enable coherent and stylistically consistent surgical report generation under limited supervision. Experiments show our approach achieves the best overall performance, with consistent gains over strong Multimodal Large Language Model (MLLM) baselines. Ablation studies further validate the effectiveness of each proposed component.

cs.CV

ReportMedSAM: Guiding Segmentation Through Radiology Reports

Free-form radiology reports contain rich clinical descriptions, yet converting them for reliable segmentation remains challenging due to the inherent variability of natural language. Existing pipelines often rely on predefined organ phrases or brittle rule-based inference-time extraction, which limits their scalability to novel anatomical structures and makes them sensitive to linguistic variations. To address this, we propose ReportMedSAM, a report-driven framework that replaces discrete extraction with a learnable concept bank. By leveraging a frozen medical vision-language encoder (BiomedCLIP), we align organ-level concept embeddings with large-scale clinical corpora through contrastive learning, establishing mutually orthogonal semantic anchors. Our approach explicitly mitigates organ-level semantic collapse and ensures high robustness against diverse clinical synonyms (e.g., "renal" vs. "kidney" ). During inference, a clinical report is embedded and matched against this concept bank to dynamically activate task-specific Mixture-of-Experts (MoE) modules. This decoupled design allows new concepts and experts to be added without retraining existing components, providing a parameter-isolated extension mechanism while keeping previously learned experts unchanged. Evaluated on the AbdomenAtlas 3.0 dataset, ReportMedSAM effectively interprets free-form reports, achieves competitive segmentation accuracy, and demonstrates seamless, non-interfering extension to novel clinical tasks.

cs.CL

OpenExtract: Automated Data Extraction for Systematic Reviews in Health

This study presents OpenExtract, an open-source pipeline for automated data extraction in large-scale systematic literature reviews. The pipeline queries large language models (LLMs) to predict data entries based on relevant sections of scientific articles. To test the efficacy of OpenExtract, we apply it to a systematic literature review in digital health and compare its outputs with those of human researchers. OpenExtract achieves precision and recall scores of > 0.8 in this task, indicating that it can be effective at extracting data automatically and efficiently. OpenExtract: https://github.com/JimAchterbergLUMC/OpenExtract.

cs.IR

SAGCNet: Spatial-Aware Graph Completion Network for Missing Slice Imputation in Population CMR Imaging

Magnetic resonance imaging (MRI) provides detailed soft-tissue characteristics that assist in disease diagnosis and screening. However, the accuracy of clinical practice is often hindered by missing or unusable slices due to various factors. Volumetric MRI synthesis methods have been developed to address this issue by imputing missing slices from available ones. The inherent 3D nature of volumetric MRI data, such as cardiac magnetic resonance (CMR), poses significant challenges for missing slice imputation approaches, including (1) the difficulty of modeling local inter-slice correlations and dependencies of volumetric slices, and (2) the limited exploration of crucial 3D spatial information and global context. In this study, to mitigate these issues, we present Spatial-Aware Graph Completion Network (SAGCNet) to overcome the dependency on complete volumetric data, featuring two main innovations: (1) a volumetric slice graph completion module that incorporates the inter-slice relationships into a graph structure, and (2) a volumetric spatial adapter component that enables our model to effectively capture and utilize various forms of 3D spatial context. Extensive experiments on cardiac MRI datasets demonstrate that SAGCNet is capable of synthesizing absent CMR slices, outperforming competitive state-of-the-art MRI synthesis methods both quantitatively and qualitatively. Notably, our model maintains superior performance even with limited slice data.

eess.IV

RefineSeg: Dual Coarse-to-Fine Learning for Medical Image Segmentation

High-quality pixel-level annotations of medical images are essential for supervised segmentation tasks, but obtaining such annotations is costly and requires medical expertise. To address this challenge, we propose a novel coarse-to-fine segmentation framework that relies entirely on coarse-level annotations, encompassing both target and complementary drawings, despite their inherent noise. The framework works by introducing transition matrices in order to model the inaccurate and incomplete regions in the coarse annotations. By jointly training on multiple sets of coarse annotations, it progressively refines the network's outputs and infers the true segmentation distribution, achieving a robust approximation of precise labels through matrix-based modeling. To validate the flexibility and effectiveness of the proposed method, we demonstrate the results on two public cardiac imaging datasets, ACDC and MSCMRseg, and further evaluate its performance on the UK Biobank dataset. Experimental results indicate that our approach surpasses the state-of-the-art weakly supervised methods and closely matches the fully supervised approach.

cs.CV

DiffKAN-Inpainting: KAN-based Diffusion model for brain tumor inpainting

Brain tumors delay the standard preprocessing workflow for further examination. Brain inpainting offers a viable, although difficult, solution for tumor tissue processing, which is necessary to improve the precision of the diagnosis and treatment. Most conventional U-Net-based generative models, however, often face challenges in capturing the complex, nonlinear latent representations inherent in brain imaging. In order to accomplish high-quality healthy brain tissue reconstruction, this work proposes DiffKAN-Inpainting, an innovative method that blends diffusion models with the Kolmogorov-Arnold Networks architecture. During the denoising process, we introduce the RePaint method and tumor information to generate images with a higher fidelity and smoother margin. Both qualitative and quantitative results demonstrate that as compared to the state-of-the-art methods, our proposed DiffKAN-Inpainting inpaints more detailed and realistic reconstructions on the BraTS dataset. The knowledge gained from ablation study provide insights for future research to balance performance with computing cost.

eess.IV

Combining multi-site Magnetic Resonance Imaging with machine learning predicts survival in paediatric brain tumours

Background Brain tumours represent the highest cause of mortality in the paediatric oncological population. Diagnosis is commonly performed with magnetic resonance imaging and spectroscopy. Survival biomarkers are challenging to identify due to the relatively low numbers of individual tumour types, especially for rare tumour types such as atypical rhabdoid tumours. Methods 69 children with biopsy-confirmed brain tumours were recruited into this study. All participants had both perfusion and diffusion weighted imaging performed at diagnosis. Data were processed using conventional methods, and a Bayesian survival analysis performed. Unsupervised and supervised machine learning were performed with the survival features, to determine novel sub-groups related to survival. Sub-group analysis was undertaken to understand differences in imaging features, which pertain to survival. Findings Survival analysis showed that a combination of diffusion and perfusion imaging were able to determine two novel sub-groups of brain tumours with different survival characteristics (p <0.01), which were subsequently classified with high accuracy (98%) by a neural network. Further analysis of high-grade tumours showed a marked difference in survival (p=0.029) between the two clusters with high risk and low risk imaging features. Interpretation This study has developed a novel model of survival for paediatric brain tumours, with an implementation ready for integration into clinical practice. Results show that tumour perfusion plays a key role in determining survival in brain tumours and should be considered as a high priority for future imaging protocols.

q-bio.QM

Distinguishing between paediatric brain tumour types using multi-parametric magnetic resonance imaging and machine learning: a multi-site study

The imaging and subsequent accurate diagnosis of paediatric brain tumours presents a radiological challenge, with magnetic resonance imaging playing a key role in providing tumour specific imaging information. Diffusion weighted and perfusion imaging are commonly used to aid the non invasive diagnosis of paediatric brain tumours, but are usually evaluated by expert qualitative review. Quantitative studies are mainly single centre and single modality. The aim of this work was to combine multi centre diffusion and perfusion imaging, with machine learning, to develop machine learning based classifiers to discriminate between three common paediatric tumour types. The results show that diffusion and perfusion weighted imaging of both the tumour and whole brain provide significant features which differ between tumour types, and that combining these features gives the optimal machine learning classifier with greater than 80 percent predictive precision. This work represents a step forward to aid in the non invasive diagnosis of paediatric brain tumours, using advanced clinical imaging.

physics.med-ph

eSource for clinical trials: Implementation and evaluation of a standards-based approach in a real world trial

Objective: The Learning Health System (LHS) requires integration of research into routine practice. eSource or embedding clinical trial functionalities into routine electronic health record (EHR) systems has long been put forward as a solution to the rising costs of research. We aimed to create and validate an eSource solution that would be readily extensible as part of a LHS. Materials and Methods: The EU FP7 TRANSFoRm project's approach is based on dual modelling, using the Clinical Research Information Model (CRIM) and the Clinical Data Integration Model of meaning (CDIM) to bridge the gap between clinical and research data structures, using the CDISC Operational Data Model (ODM) standard. Validation against GCP requirements was conducted in a clinical site, and a cluster randomised evaluation by site nested into a live clinical trial. Results: Using the form definition element of ODM, we linked precisely modelled data queries to data elements, constrained against CDIM concepts, to enable automated patient identification for specific protocols and prepopulation of electronic case report forms (e-CRF). Both control and eSource sites recruited better than expected with no significant difference. Completeness of clinical forms was significantly improved by eSource, but Patient Related Outcome Measures (PROMs) were less well completed on smartphones than paper in this population. Discussion: The TRANSFoRm approach provides an ontologically-based approach to eSource in a low-resource, heterogeneous, highly distributed environment, that allows precise prospective mapping of data elements in the EHR. Conclusion: Further studies using this approach to CDISC should optimise the delivery of PROMS, whilst building a sustainable infrastructure for eSource with research networks, trials units and EHR vendors.

cs.CY

The What, Who, Where, When, Why and How of Context-Awareness

The understanding of context and context-awareness is very important for the areas of handheld and ubiquitous computing. Unfortunately, at present, there has not been a satisfactory definition of these two concepts that would lead to a more effective communication in humancomputer interaction. As a result, on the one hand, application designers are not able to choose what context to use in their applications and on the other, they cannot determine the type of context-awareness behaviours their applications should exhibit. In this work, we aim to provide answers to some fundamental questions that could enlighten us on the definition of context and its functionality.

cs.HC