SearcharxivSearch

arXiv subjects

Xiaofeng Yang

Publications and source records attributed to Xiaofeng Yang.

At least 19 recordsLinked to original sources

A radiographic world model for clinical reasoning and evidence generation

Medical imaging artificial intelligence (AI) is commonly developed as separate mappings from radiographs to diagnostic outputs or from clinical descriptions to generated images, although both arise from the same underlying radiographic state. A world-model formulation instead seeks to learn an internal representation of this state that can support both clinical readout and conditional simulation of radiographic observations. Here we introduce MedDream, a radiographic world model that learns a shared continuous latent state from paired chest radiograph-text observations for diagnostic reasoning and report-conditioned evidence generation. MedDream was pretrained on 2.65 million leakage-controlled chest radiograph-text pairs curated from 4.40 million candidates. Across eight clinical datasets and two independent reader cohorts, MedDream outperformed leading diagnostic and generative comparators. For diagnostic reasoning, MedDream showed strong generalization across disease recognition, label-scarce adaptation, severity assessment, and localization, while MedDream-supported review increased mean resident concordance with independent radiologist consensus from 56.3% to 63.0%. For evidence generation, MedDream produced radiographs that preserved clinically relevant pathology and improved downstream performance on held-out real data, with synthetic augmentation increasing external VinDr-CXR macro-AUROC from 76.4% to 81.4%. More importantly, conditioning generation on prespecified subgroup performance gaps enabled targeted evidence construction, increasing weighted F1 by 3.1 percentage points in Asian patients, whereas matched-volume unguided augmentation decreased it by 2.3 points. These findings establish radiographic world models as a path toward medical AI that learns clinically meaningful internal states for interpreting, simulating, and constructing evidence for clinical use.

cs.AI

Studies on the dark sector interaction from joint analysis of cosmological probes

We test whether constraints on the nonlinear interaction $\xi$IDE are stable under different treatments of the Type Ia supernovae absolute calibration. \textit{Fermi} GRBs measurements and the Amati-relation parameters are fitted jointly with PantheonPlus SNe Ia, DESI DR2 BAO, and an updated cosmic-chronometer compilation. We compare the PantheonPlus-SH0ES route, which retains the SN absolute calibration, with the PantheonPlus-only route, in which the SN absolute magnitude is analytically marginalized. The GOLD GRB sample is adopted for the main analysis, while the FULL GRB sample is used to assess sample dependence. For the interaction parameter $\gamma\equiv\xi+3w$, where $\gamma=0$ denotes the non-interacting limit, the GOLD sample gives $\gamma=1.453^{+1.297}_{-1.597}$ for the PantheonPlus-SH0ES and $\gamma=-0.634^{+1.668}_{-2.486}$ for the PantheonPlus-only. Although the posterior medians correspond to opposite directions of energy transfer, neither route excludes $\gamma=0$ at 68\% credibility, and the reconstructed interaction rate remains consistent with zero over the redshift range considered. Replacing the GOLD sample with the FULL sample produces negligible changes in the interaction constraints. Moreover, $w$CDM and CPL achieve likelihood improvements comparable to that of $\xi$IDE, while the information criteria do not consistently favor the interacting model. A redshift-bin diagnostic finds no significant redshift evolution of the Amati relation. We find no compelling evidence for a dark sector interaction that is robust to the choice of SN calibration or specifically favored over noninteracting dark energy extensions.

astro-ph.CO

Motion Artifact-Aware Self-Supervised Representation Learning for 3D Brain MRI Motion Artifact Reduction

Patient motion remains a source of image degradation in brain MRI, leading to signal loss, blurring, and geometric distortion that compromise quantitative analysis. Existing deep learning methods for motion correction typically rely on paired clean-corrupted data or k-space acquisitions, which are rarely available in clinical settings. We propose SSRL-MAR, a motion artifact-aware unpaired representation learning framework for motion artifact reduction that requires neither paired training data nor explicit motion labels. SSRL-MAR employed a three-stage training strategy: (1) contrastive learning on 3D patches to extract motion representations by contrasting clean and synthetically corrupted images, (2) a motion artifact-aware synthesis network to generate motion artifacts from clean scans, and (3) a motion artifact-aware generator to restore clean volumes using the learned degrader for self-supervised supervision. On in-silico dataset, SSRL-MAR achieved PSNR 23.81dB, SSIM 91.55%, and NMSE 0.79%. On in-vivo MR-ART dataset, the pretrained model reduced motion distortion, and unsupervised domain adaptation further improved anatomical fidelity. Against a source-only supervised model trained on the same simulated pairs, SSRL-MAR improved PSNR by up to 2.0 dB on MR-ART after unsupervised domain adaptation, and remained within 0.25-0.47 dB of an oracle supervised model that requires real paired data unavailable in practice. At the milder motion level, volumetric error in structures such as the corpus callosum and ventricular system decreased by more than 50%, confirming improved neuroanatomical consistency. These results indicate that SSRL-MAR provides a robust and scalable image-domain solution for 3D brain MRI motion correction, enabling reliable structural quantification in large-scale neuroimaging studies without requiring prospectively acquired pairs or acquisition-specific calibration.

cs.CV

MRI super-resolution in ten sampling steps using a diffusion bridge model

Objective. MRI provides excellent soft-tissue contrast, but long acquisition times can cause patient discomfort and lead to motion artifacts, forcing a trade-off between spatial resolution and scan time. Diffusion-based super-resolution (SR) reconstructs high-resolution (HR) images from low-resolution (LR) inputs, but typically needs many sampling steps and initializes from a Gaussian prior ill-suited to image restoration. We developed an efficient diffusion framework that reconstructs HR MRI directly from LR data. Approach. We propose super-resolution diffusion bridge model (SR-DBM), a super-resolution diffusion bridge model that casts SR as a stochastic transport between the LR and HR image distributions. Through a Doob's h-transform of a mean-reverting stochastic differential equation, SR-DBM pins the process to the paired HR and LR images at its endpoints, initializing reconstruction from the measured anatomy rather than from Gaussian noise. The HR image is recovered by a deterministic reverse trajectory in which a network predicts the clean image at each of only ten sampling steps. We evaluated SR-DBM on ultra-high-field 7T brain T1 MP2RAGE maps and pelvic T2-weighted prostate images against nine comparison methods using PSNR, SSIM, GMSD, and LPIPS. Main results. SR-DBM attained the highest PSNR and SSIM and the lowest GMSD on both datasets (brain: 27.66+-1.52 dB, 0.96+-0.02, 7.96+-1.86$; prostate: 27.87+-2.29 dB, 0.80+-0.05, 8.38+- 1.44), with statistically significant gains over every comparison method (two-sided Wilcoxon signed-rank test with Holm correction, p<0.05). The strongest baseline, SR-EMamba, ranked second. Qualitatively, SR-DBM produced the smallest residual errors and best preserved fine structures and lesions.

cs.CV

Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model

Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is time-consuming. Models such as nnU-Net may generalize imperfectly and lack clinician-directed text correction. Purpose: We investigated adapting a three-dimensional (3D) vision-language foundation model for text-guided brain tumor segmentation refinement. Methods: We developed a lightweight VoxTell-based framework. Pretrained VoxTell generated initial masks. Oracle prompts derived from segmentation errors encoded target, action, location, imaging evidence, edit size, and preservation constraints. Frozen Qwen/VoxTell prompt embeddings were injected through trainable projections into its multiscale decoder conditioning; other weights remained frozen. Training, validation, and testing used 901, 100, and 250 BraTS-GLI cases. Cross-dataset transfer was evaluated on 100 meningioma, metastasis, pediatric tumor, and UPENN-GBM cases. Results: On the internal test set using post-contrast T1-weighted input, correct instructions improved subregion Dice similarity coefficient (DSC; enhancing tumor, edema, and necrotic/non-enhancing core) from $0.774\pm0.158$ to $0.796\pm0.137$. They outperformed blank prompts ($0.762\pm0.155$; Holm-adjusted $p<0.001$, $d_z=0.71$) and contradictory prompts ($0.770\pm0.163$; $p<0.001$, $d_z=0.48$). In cross-dataset testing, correct instructions improved DSC from $0.527\pm0.287$ to $0.550\pm0.278$ and outperformed contradictory instructions ($0.504\pm0.275$; $p<0.001$, $d_z=0.43$). Conclusion: A 3D vision-language foundation model can perform instruction-guided refinement of glioma subregion segmentations. Sensitivity to correct, blank, and contradictory prompts suggests text-dependent contour editing rather than nonspecific post-processing, supporting further evaluation as a clinician-in-the-loop tool.

cs.CV

Anticipatory Digital Twins for Online Head-and-Neck Adaptive Proton Therapy via Foundation-Model Registration

Head-and-neck (HN) proton therapy is highly sensitive to anatomical change over a 4-to-6-week course, as tumor shrinkage, weight loss, and setup variation can misposition the Bragg peak near critical organs such as the parotids, oral cavity, brainstem, and spinal cord, leading to target underdosing or organ-at-risk overdosing. Online adaptive proton therapy replans on the anatomy of the day, yet standard workflows rely on offline replanning that requires repeated CT acquisition and roughly a week of preparation, adding burden, cost, and delay. We investigate whether a patient's treatment-day anatomy can be predicted before image acquisition by transferring longitudinal change from a population database. We propose a digital-twin framework built on a pretrained foundation-model deformable registration network used without patient-specific training. A first registration aligns a prior patient's planning CT to the target and carries the prior's during-treatment quality assurance CT (QACT) into the target frame; a second registration estimates the prior's planning-to-QACT change, which is then applied to the target's own planning CT to synthesize predicted CTs (pdCTs) with propagated contours. Using 88 HN patients, each with a planning CT and three QACTs, we show that pdCTs better match treatment-day anatomy than the static planning CT. Compared with the planning CT alone, normalized cross-correlation improves by 22.8%, Dice for organs-at-risk by 20.2%, and CT-number error decreases by 23.4%. Gains are largest for patients with major anatomical change and negligible when anatomy is stable. This cross-patient motion transfer leverages the digital-twin concept to anticipate treatment-day anatomy, enabling personalized online adaptive proton therapy without repeated imaging.

physics.med-ph

Circulating Lymphocytes Preservation in Lung Cancer Stereotactic Body Radiation Therapy with Ultra-Fast Proton Delivery Using Modularized Pin Ridge Filters

Purpose: Radiation-induced lymphopenia is an increasingly recognized toxicity in lung radiotherapy and has been linked to radiation exposure to circulating lymphocytes (CL). In intensity-modulated proton therapy (IMPT), prolonged pencil beam scanning (PBS) delivery may increase CL dose. We recently developed a patient-specific pin ridge filter (pRF) framework that enables ultra-fast proton delivery with a single beam energy. This study evaluated whether pRF-based lung stereotactic body radiotherapy (SBRT) plans delivered at conventional (pRFCONV) and FLASH dose rates (pRFFLASH) improve immune sparing using time-resolved blood dose accumulation and CL survival modeling. Methods: pRF plans were created for 10 lung SBRT patients previously treated with IMPT. PBS delivery simulations modeled spot delivery, scanning, and energy switching. Blood dose-volume histograms (bDVHs) were calculated with the hematological dose framework. CL survival fractions (SF) were estimated from bDVHs with saturation and linear-quadratic models derived from in-vitro survival data for CD4/CD8 CL. Results: Compared with IMPT, pRFCONV/pRFFLASH plans reduced delivery time (mean reductions: 85.3/99.9%) and irradiated blood volume per fraction (mean reductions: 52.9/81.3%). pRFCONV/pRFFLASH plans reduced blood V5cGy by 26.4/39.4%, and V50cGy by 4.5/6.9%, respectively. pRF plans improved modeled CL survival across all models and subpopulations. Unstimulated CD4/CD8 CL had the largest SF differences, for which mean saturation-model SF improved by 7.9/8.6% for pRFCONV (p=0.03/0.02) and 9.6/10.4% for pRFFLASH (p=0.02/0.01), respectively. Conclusion: pRF plans improved modeled CL survival by significantly shortening delivery time and reducing irradiation of circulating blood. Our findings suggest that pRF's ultra-fast delivery may provide a practical strategy for immune sparing in proton lung SBRT.

physics.med-ph

One-for-All Adaptive Radiotherapy Planning Agent: A Foundation Framework for Daily CBCT-guided Radiotherapy

In this work, we introduce the One-for-All Adaptive Radiotherapy Planning Agent, a unified foundation-model-based system that performs complete, treatment-specific online adaptive planning directly from daily cone-beam CT in under two minutes. The agent first autonomously predicts all essential planning components, including synthetic CT generation, multimodal alignment, and tumor/organ segmentation. It then intelligently leverages these outputs to execute the final clinical plan design, providing a comprehensive, automated solution for daily treatment. We also demonstrate that the agent enables clinicians to define planning with intent and intervene at critical decision points, ensuring a "human-in-the-loop" framework that generates acceptable plans before final approval. Evaluated on multiple datasets spanning head-and-neck, lung, abdominal, and prostate cancers with both photon and proton therapy, the proposed framework achieves clinically acceptable accuracy and plan quality comparable to clinically generated treatment plans, with target dose errors (D98) generally within 2.0 Gy of the reference plan. The strong performance of the One-for-All agent highlights the promise of a unified foundation-model approach and opens opportunities for fast, scalable, and fully automated online adaptive radiotherapy across diverse clinical scenarios.

physics.med-ph

Artificial Intelligence Across the Cardiac Amyloidosis Diagnostic Pathway: From Single-Modality Detection to Multimodal Clinical Integration

Cardiac amyloidosis (CA) is increasingly recognized but remains substantially underdiagnosed, because its clinical and imaging phenotype overlaps with more common cardiomyopathies. Definitive subtype assignment and management further require integration of multimodal evidence to distinguish transthyretin from light chain disease. Machine learning and deep learning have been applied across the diagnostic and management pathway. These applications span ECG, echocardiography, and health record-based case finding, as well as CMR and nuclear interpretation, including SPECT/CT biomarker quantification, prognostic modeling, and treatment response assessment. This narrative review synthesizes these studies by clinical tasks, namely screening, detection, quantification, prognosis, and treatment response monitoring, rather than by input modality. This task-based organization clarifies why apparently similar AI models require different cohorts, reference standards, evaluation metrics, and implementation thresholds. The evidence reveals a maturity gradient. Binary detection and AI assisted quantification on bone scintigraphy and SPECT/CT are closest to clinical translation. Detection is supported by large externally validated cohorts, and quantification by interpretable, outcome linked measurement of myocardial tracer burden. By contrast, subtype aware classification, prognostic risk stratification, and treatment response monitoring remain at an early stage. These tasks are limited by small cohorts, enriched retrospective designs, heterogeneous labels, incomplete external validation, and uncertain calibration in realistic prevalence settings. Across tasks, high discrimination alone is insufficient.

physics.med-ph

EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning

Local editing of 3D objects remains a long-standing challenge. When interacting with 3D content, humans naturally tend to specify a coarse region of interest for modification rather than defining precise editing boundaries. However, previous methods rely on fully edited 2D images, precise 3D masks, or redundant pipelines, which present a gap. To bridge this gap, we propose EditVerse3D, a novel 3D editing framework that enables high-quality object editing under such coarse guidance. Our approach takes as input a 3D object to be edited, a coarse 3D bounding box indicating the target region, and a reference 2D image describing the desired modification. It produces a coherent, high-fidelity edited 3D object. To facilitate this editing, we introduce a novel region-aware adaptive loss that emphasizes hard-to-learn regions and balances the objective between target and preserved areas. Complementing our loss function, we enhance model robustness and generalization through targeted data augmentations, such as training with scaled 3D masks and filtering out unrealistic editing pairs. We construct a large-scale 3D editing dataset derived from parts information. Extensive experiments demonstrate that EditVerse3D achieves superior visual quality and quantitative performance compared to existing 3D editing approaches. Please visit our project page at https://editverse3d.github.io.

cs.CV

Interacting dark energy constraints from Fermi GRBs and Pantheon+ SNe Ia with full GRB covariance

The standard $\Lambda$CDM model faces long-standing theoretical and observational problems, such as the Hubble tension, which motivate extensions beyond $\Lambda$CDM, including interacting dark energy (IDE). Type Ia supernovae (SNe Ia) are precise probes of the late-time expansion history, while gamma-ray bursts (GRBs) can extend distance measurements to higher redshifts. However, GRB cosmology depends on the calibration of luminosity relations, the covariance treatment, and the intrinsic scatter. In this work, we use 15 years of Fermi/GBM long-GRB observations and Pantheon+ SNe Ia to test whether current distance data provide evidence in favor of IDE models over $\Lambda$CDM. We compare four flat models: $\Lambda$CDM, $w$CDM, IDE-$\rho_{\rm de}$, and IDE-$\rho_{\rm c}$. The GRB covariance is constructed by propagating the Amati-relation calibration covariance, and the GRB intrinsic scatter is sampled as a nuisance parameter. A diagonal GRB covariance is also considered as a robustness test. With the full GRB covariance, both the GOLD and FULL samples give $H_0\simeq 72.8~{\rm km~s^{-1}~Mpc^{-1}}$ in $\Lambda$CDM. The IDE models do not improve the fit enough to compensate for their extra parameters, and the BIC favors the simpler $\Lambda$CDM model. The diagonal-covariance test gives the same model-selection conclusion, although it changes the fitted GRB intrinsic scatter. We conclude that, for the two interaction forms considered here and at the present level of GRB systematics, current GRB and Pantheon+ data do not provide evidence for interacting dark energy. Current GRBs mainly provide a high-redshift extension of the Hubble diagram and test the shape of the expansion history.

astro-ph.CO

Testing the cosmological principle with quasars

The inferred velocity is consistent at the 1.56 {\sigma} level with the value of 370 km/s from a purely kinematic interpretation of the CMB dipole. Based on the motion direction component analysis, we have not found any significant deviation from cosmological principle in current released quasars data. The cosmological principle posits that the universe is homogeneous and isotropic on the large scales. In history, the cosmological principle was confirmed by various cosmological observations from CMB to large scale structure. However, several new challenges to the cosmological principle were reported in recent years, particularly in radio observations from overdispersed radio source counts to quasars. Here, we firstly present studies on the peculiar velocity of large-scale anisotropy by measuring the dipole signal from the DESI DR1 catalogue with a sample of 1,176,570 quasars (0.8 < z < 3.0). Our analysis reveals the peculiar velocity of $|v| = 443.8 \pm 204.1$ km/s towards $(l, b) = (107.4^\circ \pm 86.8^\circ, 28.4^\circ \pm 45.2^\circ)$ in Galactic coordinates.The motion direction deviates from the CMB dipole (264.02$^\circ$, 48.253$^\circ$). The inferred velocity is consistent at the 1.56 $\sigma$ level with the value of 370 km/s from a purely kinematic interpretation of the CMB dipole. Based on the motion direction component analysis, we have not found any significant deviation from cosmological principle in current released quasars data.

astro-ph.CO

Universal CT Representations from Anatomy to Disease Phenotype through Agglomerative Pretraining

Computed tomography (CT) is a central to three-dimensional medical imaging, yet CT-based artificial intelligence remains fragmented across task-specific models for segmentation, classification, registration, and report analysis. Here we present FlexiCT, a family of CT foundation models trained by agglomerative continual pretraining on 266,227 CT volumes from 56 publicly available datasets, forming a large-scale public resource for CT representation learning. FlexiCT uses agglomerative pretraining across three stages: two-dimensional axial pretraining, three-dimensional anatomical pretraining and report-guided semantic alignment. This training strategy supports slice-level, volume-level and vision-language analysis. Across five downstream task families (segmentation, classification, registration, vision-language understanding and clinical retrieval), FlexiCT matches or exceeds prior task-specific approaches on multiple benchmarks. Its embeddings further organize CT scans along gradients associated with various tumor stages, suggesting that CT foundation models can capture imaging features relevant to disease phenotype characterization. Project page and code are available at: https://ricklisz.github.io/flexict.github.io and https://github.com/ricklisz/FlexiCT.

cs.CV

BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning

Brain MRI underpins a wide range of neuroscientific and clinical applications, yet most learning-based methods remain task-specific and require substantial labeled data. Here we show that a single self-supervised representation can generalize across heterogeneous brain MRI endpoints. We trained BrainDINO, a self-distilled foundation model, on approximately 6.6 million unlabeled axial slices from 20 datasets encompassing broad variation in population, disease, and acquisition setting. Using a frozen encoder with lightweight task heads, BrainDINO supported transfer across tumor segmentation, neurodegenerative and neurodevelopmental conditions classification, brain age estimation, post-stroke temporal prediction, molecular status prediction, MRI sequence classification, and survival modeling. Across tasks and supervision regimes, BrainDINO consistently equaled or exceeded natural-image and MRI-specific self-supervised baselines, with particularly strong advantages under label scarcity. Representation analyses further showed anatomically organized and pathology-sensitive feature structure in the absence of task-specific supervision. Our findings indicate that large-scale slice-wise self-supervised learning can yield a unified brain MRI representation that supports diverse neuroimaging tasks without volumetric pretraining or full-network fine-tuning, establishing a scalable foundation for robust and data-efficient brain imaging analysis. Code is available at https://github.com/mclwu22/BrainDINO

cs.LG

Measurement of the Weyl Potential Evolution and $E_G$ Statistic from KiDS-1000, BOSS and 2dFLenS

A recently developed model-independent approach to measuring the Weyl potential has shown some tensions with $\Lambda$CDM (Lambda Cold Dark Matter) in DES (Dark Energy Survey) Y3 data. We apply this framework to Kilo-Degree Survey (KiDS-1000) weak lensing and BOSS/2dFLenS galaxy clustering using the KiDS Cosmology Analysis Pipeline (KCAP) in two redshift bins. To apply this method to the BOSS spectroscopic wedges, we take the early-time power spectrum from a redshift where General Relativity (GR) is recovered and jointly fit the late-time observables $\hat{b}$, $\hat{f}$, and $\hat{J}$. With Planck18 priors, the CMASS measurement of $\hat{J}$ is $1.50\sigma$ below the GR prediction, while the LOWZ measurement remains consistent with it. The decrease in the central value of $\hat{J}$ from LOWZ to CMASS provides a mild indication of the previously reported ``Lensing is low'' behavior. Without Planck18 priors, we obtain $E_G(0.38)=0.503\pm0.054$ and $E_G(0.61)=0.444\pm0.054$, both consistent with GR. Furthermore, phenomenological gravity tests also show no significant preference for modified gravity.

astro-ph.CO

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language. This paradigm is fundamentally limited in clinical scenarios, where accurate answers often depend on subtle, localized visual evidence that cannot be reliably preserved in static embeddings. We propose \textsc{MedLVR}, a latent visual reasoning framework that introduces an explicit visual evidence state into autoregressive decoding. Instead of relying solely on text-based intermediate reasoning, \textsc{MedLVR} interleaves a short latent reasoning segment within the decoder by reusing hidden states as continuous latent steps, enabling iterative preservation and refinement of query-relevant visual evidence before answer generation. To support effective visual supervision, we adopt a two-stage training strategy: region of interest (ROI)-supervised fine-tuning aligns latent states with clinically relevant image evidence, and Visual-Latent Policy Optimization (VLPO) further optimizes latent reasoning and answer generation under outcome-level rewards. Experiments on OmniMedVQA and five external medical VQA benchmarks show that \textsc{MedLVR} consistently outperforms recent reasoning baselines and improves the average score over the Qwen2.5-VL-7B backbone from 48.3\% to 53.4\%. These results show that latent visual reasoning provides an effective mechanism for preserving diagnostically relevant visual evidence and improving the reliability of medical VQA.

cs.CV

Distilling Photon-Counting CT into Routine Chest CT through Clinically Validated Degradation Modeling

Photon-counting CT (PCCT) provides superior image quality with higher spatial resolution and lower noise compared to conventional energy-integrating CT (EICT), but its limited clinical availability restricts large-scale research and clinical deployment. To bridge this gap, we propose SUMI, a simulated degradation-to-enhancement method that learns to reverse realistic acquisition artifacts in low-quality EICT by leveraging high-quality PCCT as reference. Our central insight is to explicitly model realistic acquisition degradations, transforming PCCT into clinically plausible lower-quality counterparts and learning to invert this process. The simulated degradations were validated for clinical realism by board-certified radiologists, enabling faithful supervision without requiring paired acquisitions at scale. As outcomes of this technical contribution, we: (1) train a latent diffusion model on 1,046 PCCTs, using an autoencoder first pre-trained on both these PCCTs and 405,379 EICTs from 145 hospitals to extract general CT latent features that we release for reuse in other generative medical imaging tasks; (2) construct a large-scale dataset of over 17,316 publicly available EICTs enhanced to PCCT-like quality, with radiologist-validated voxel-wise annotations of airway trees, arteries, veins, lungs, and lobes; and (3) demonstrate substantial improvements: across external data, SUMI outperforms state-of-the-art image translation methods by 15% in SSIM and 20% in PSNR, improves radiologist-rated clinical utility in reader studies, and enhances downstream top-ranking lesion detection performance, increasing sensitivity by up to 15% and F1 score by up to 10%. Our results suggest that emerging imaging advances can be systematically distilled into routine EICT using limited high-quality scans as reference.

cs.CV

CBCT-Based Synthetic CT Generation Using Conditional Flow Matching Model

Daily or weekly cone-beam computed tomography (CBCT) is employed in image-guided radiotherapy (IGRT) for precise patient alignment. However, its clinical utility in quantitative tasks is hindered by severe artifacts and inaccurate Hounsfeld unit (HU). It is essential to enhance CBCT image quality to a level comparable with that of conventional CT scans. This study proposed a conditional flow matching model that gradually transforms a sample from normal distribution to the corresponding CT sample conditioned on the input CBCT image. The proposed model was trained using CBCT and deformed planning CT (dpCT) image pairs in a supervised learning scheme. The feasibility of the conditional flow matching model was verified using studies of brain, head-and-neck (HN), and lung patients. The quantitative performance was evaluated using three metrics, including mean absolute error (MAE), peak signal-to-noise ratio (PSNR), and normalized cross-correlation (NCC). The proposed flow matching model was also compared to other flow matching and diffusion-based generative models for sCT generation. The proposed flow matching model effectively reduced multiple types of artifacts on CBCT images in all the studies. In the study of brain patient, the MAE, PSNR, and NCC of the sCT were improved to 26.02 HU, 32.35 dB, and 0.99, respectively, from 40.63 HU, 27.87 dB, and 0.98 on the CBCT images. In the study of HN patient, the metrics were improved to 33.17 HU, 28.68 dB, 0.98 from 38.99 HU, 27.00 dB, 0.98. In the lung patient study, the metrics were 25.09 HU, 32.81 dB, 0.99 and 32.90 HU, 30.48 dB, 0.98 for sCT and CBCT, respectively. The proposed conditional flow matching model effectively synthesizes high-quality CT-like images from CBCT, achieving accurate HU representation and artifact reduction. This enables more reliable organ segmentation and dose calculation in CBCT-guided online ART workflows.

physics.med-ph