Searcharxiv⌕ Search

arXiv subjects

William W. Wong

Publications and source records attributed to William W. Wong.

9 recordsLinked to original sources

A Multimodal Clinically Informed Coarse-to-Fine Framework for Longitudinal CT Registration in Proton Therapy

Proton therapy offers superior organ-at-risk sparing but is highly sensitive to anatomical changes, making accurate deformable image registration (DIR) across longitudinal CT scans essential. Conventional DIR methods are often too slow for emerging online adaptive workflows, while existing deep learning-based approaches are primarily designed for generic benchmarks and underutilize clinically relevant information beyond images. To address this gap, we propose a clinically scalable coarse-to-fine deformable registration framework that integrates multimodal information from the proton radiotherapy workflow to accommodate diverse clinical scenarios. The model employs dual CNN-based encoders for hierarchical feature extraction and a transformer-based decoder to progressively refine deformation fields. Beyond CT intensities, clinically critical priors, including target and organ-at-risk contours, dose distributions, and treatment planning text, are incorporated through anatomy- and risk-guided attention, text-conditioned feature modulation, and foreground-aware optimization, enabling anatomically focused and clinically informed deformation estimation. We evaluate the proposed framework on a large-scale proton therapy DIR dataset comprising 1,222 paired planning and repeat CT scans across multiple anatomical regions and disease types. Extensive experiments demonstrate consistent improvements over state-of-the-art methods, enabling fast and robust clinically meaningful registration.

cs.CV↗

An Automated Retrieval-Augmented Generation LLaMA-4 109B-based System for Evaluating Radiotherapy Treatment Plans

Purpose: To develop a retrieval-augmented generation (RAG) system powered by LLaMA-4 109B for automated, protocol-aware, and interpretable evaluation of radiotherapy treatment plans. Methods and Materials: We curated a multi-protocol dataset of 614 radiotherapy plans across four disease sites and constructed a knowledge base containing normalized dose metrics and protocol-defined constraints. The RAG system integrates three core modules: a retrieval engine optimized across five SentenceTransformer backbones, a percentile prediction component based on cohort similarity, and a clinical constraint checker. These tools are directed by a large language model (LLM) using a multi-step prompt-driven reasoning pipeline to produce concise, grounded evaluations. Results: Retrieval hyperparameters were optimized using Gaussian Process on a scalarized loss function combining root mean squared error (RMSE), mean absolute error (MAE), and clinically motivated accuracy thresholds. The best configuration, based on all-MiniLM-L6-v2, achieved perfect nearest-neighbor accuracy within a 5-percentile-point margin and a sub-2pt MAE. When tested end-to-end, the RAG system achieved 100% agreement with the computed values by standalone retrieval and constraint-checking modules on both percentile estimates and constraint identification, confirming reliable execution of all retrieval, prediction and checking steps. Conclusion: Our findings highlight the feasibility of combining structured population-based scoring with modular tool-augmented reasoning for transparent, scalable plan evaluation in radiation therapy. The system offers traceable outputs, minimizes hallucination, and demonstrates robustness across protocols. Future directions include clinician-led validation, and improved domain-adapted retrieval models to enhance real-world integration.

cs.AI↗

Diffusion Transformer-based Universal Dose Denoising for Pencil Beam Scanning Proton Therapy

Purpose: Intensity-modulated proton therapy (IMPT) offers precise tumor coverage while sparing organs at risk (OARs) in head and neck (H&N) cancer. However, its sensitivity to anatomical changes requires frequent adaptation through online adaptive radiation therapy (oART), which depends on fast, accurate dose calculation via Monte Carlo (MC) simulations. Reducing particle count accelerates MC but degrades accuracy. To address this, denoising low-statistics MC dose maps is proposed to enable fast, high-quality dose generation. Methods: We developed a diffusion transformer-based denoising framework. IMPT plans and 3D CT images from 80 H&N patients were used to generate noisy and high-statistics dose maps using MCsquare (1 min and 10 min per plan, respectively). Data were standardized into uniform chunks with zero-padding, normalized, and transformed into quasi-Gaussian distributions. Testing was done on 10 H&N, 10 lung, 10 breast, and 10 prostate cancer cases, preprocessed identically. The model was trained with noisy dose maps and CT images as input and high-statistics dose maps as ground truth, using a combined loss of mean square error (MSE), residual loss, and regional MAE (focusing on top/bottom 10% dose voxels). Performance was assessed via MAE, 3D Gamma passing rate, and DVH indices. Results: The model achieved MAEs of 0.195 (H&N), 0.120 (lung), 0.172 (breast), and 0.376 Gy[RBE] (prostate). 3D Gamma passing rates exceeded 92% (3%/2mm) across all sites. DVH indices for clinical target volumes (CTVs) and OARs closely matched the ground truth. Conclusion: A diffusion transformer-based denoising framework was developed and, though trained only on H&N data, generalizes well across multiple disease sites.

physics.med-ph↗

Noisy probing dose facilitated dose prediction for pencil beam scanning proton therapy: physics enhances generalizability

Purpose: Prior AI-based dose prediction studies in photon and proton therapy often neglect underlying physics, limiting their generalizability to handle outlier clinical cases, especially for pencil beam scanning proton therapy (PBSPT). Our aim is to design a physics-aware and generalizable AI-based PBSPT dose prediction method that has the underlying physics considered to achieve high generalizability to properly handle the outlier clinical cases. Methods and Materials: This study analyzed PBSPT plans of 103 prostate and 78 lung cancer patients from our institution,with each case comprising CT images, structure sets, and plan doses from our Monte-Carlo dose engine (serving as the ground truth). Three methods were evaluated in the ablation study: the ROI-based method, the beam mask and sliding window method, and the noisy probing dose method. Twelve cases with uncommon beam angles or prescription doses tested the methods' generalizability to rare treatment planning scenarios. Performance evaluation used DVH indices, 3D Gamma passing rates (3%/2mm/10%), and dice coefficients for dose agreement. Results: The noisy probing dose method showed improved agreement of DVH indices, 3D Gamma passing rates, and dice coefficients compared to the conventional methods for the testing cases. The noisy probing dose method showed better generalizability in the 6 outlier cases than the ROI-based and beam mask-based methods with 3D Gamma passing rates (for prostate cancer, targets: 89.32%$\pm$1.45% vs. 93.48%$\pm$1.51% vs. 96.79%$\pm$0.83%, OARs: 85.87%$\pm$1.73% vs. 91.15%$\pm$1.13% vs. 94.29%$\pm$1.01%). The dose predictions were completed within 0.3 seconds. Conclusions: We've devised a novel noisy probing dose method for PBSPT dose prediction in prostate and lung cancer patients. With more physics included, it enhances the generalizability of dose prediction in handling outlier clinical cases.

physics.med-ph↗

Artificial Intelligence-Facilitated Online Adaptive Proton Therapy Using Pencil Beam Scanning Proton Therapy

We propose an oAPT workflow that incorporates all these functionalities and validate its clinical implementation feasibility with prostate patients. AI-based auto-segmentation tool AccuContourTM (Manteia, Xiamen, China) was seamlessly integrated into oAPT. Initial spot arrangement tool on the vCT for re-optimization was implemented using raytracing. An LET-based biological effect evaluation tool was developed to assess the overlap region of high dose and high LET in selected OARs. Eleven prostate cancer patients were retrospectively selected to verify the efficacy and efficiency of the proposed oAPT workflow. The time cost of each component in the workflow was recorded for analysis. The verification plan showed significant degradation of the CTV coverage and rectum and bladder sparing due to the interfractional anatomical changes. Re-optimization on the vCT resulted in great improvement of the plan quality. No overlap regions of high dose and high LET distributions were observed in bladder or rectum in re-plans. 3D Gamma analyses in PSQA confirmed the accuracy of the re-plan doses before delivery (Gamma passing rate = 99.57%), and after delivery (98.59%). The robustness of the re-plans passed all clinical requirements. The average time for the complete execution of the workflow was 9.12minutes, excluding manual intervention time. The AI-facilitated oAPT workflow was demonstrated to be both efficient and effective by generating a re-plan that significantly improved the plan quality in prostate cancer treated with PBSPT.

physics.med-ph↗

Benchmarking a foundation LLM on its ability to re-label structure names in accordance with the AAPM TG-263 report

Purpose: To introduce the concept of using large language models (LLMs) to re-label structure names in accordance with the American Association of Physicists in Medicine (AAPM) Task Group (TG)-263 standard, and to establish a benchmark for future studies to reference. Methods and Materials: The Generative Pre-trained Transformer (GPT)-4 application programming interface (API) was implemented as a Digital Imaging and Communications in Medicine (DICOM) storage server, which upon receiving a structure set DICOM file, prompts GPT-4 to re-label the structure names of both target volumes and normal tissues according to the AAPM TG-263. Three disease sites, prostate, head and neck, and thorax were selected for evaluation. For each disease site category, 150 patients were randomly selected for manually tuning the instructions prompt (in batches of 50) and 50 patients were randomly selected for evaluation. Structure names that were considered were those that were most likely to be relevant for studies utilizing structure contours for many patients. Results: The overall re-labeling accuracy of both target volumes and normal tissues for prostate, head and neck, and thorax cases was 96.0%, 98.5%, and 96.9% respectively. Re-labeling of target volumes was less accurate on average except for prostate - 100%, 93.1%, and 91.1% respectively. Conclusions: Given the accuracy of GPT-4 in re-labeling structure names of both target volumes and normal tissues as presented in this work, LLMs are poised to be the preferred method for standardizing structure names in radiation oncology, especially considering the rapid advancements in LLM capabilities that are likely to continue.

physics.med-ph↗

Modelling small block aperture in an in-house developed GPU-accelerated Monte Carlo-based dose engine for pencil beam scanning proton therapy

Purpose: To enhance an in-house graphic-processing-unit (GPU) accelerated virtual particle (VP)-based Monte Carlo (MC) proton dose engine (VPMC) to model aperture blocks in both dose calculation and optimization for pencil beam scanning proton therapy (PBSPT)-based stereotactic radiosurgery (SRS). Methods and Materials: A block aperture module was integrated into VPMC. VPMC was validated by an opensource code, MCsquare, in eight water phantom simulations with 3cm thick brass apertures: four were with aperture openings of 1, 2, 3, and 4cm without a range shifter, while the other four were with same aperture opening configurations with a range shifter of 45mm water equivalent thickness. VPMC was benchmarked with MCsquare and RayStation MC for 10 patients with small targets (average volume 8.4 cc). Finally, 3 patients were selected for robust optimization with aperture blocks using VPMC. Results: In the water phantoms, 3D gamma passing rate (2%/2mm/10%) between VPMC and MCsquare were 99.71$\pm$0.23%. In the patient geometries, 3D gamma passing rates (3%/2mm/10%) between VPMC/MCsquare and RayStation MC were 97.79$\pm$2.21%/97.78$\pm$1.97%, respectively. The calculation time was greatly decreased from 112.45$\pm$114.08 seconds (MCsquare) to 8.20$\pm$6.42 seconds (VPMC), both having statistical uncertainties of about 0.5%. The robustly optimized plans met all the dose-volume-constraints (DVCs) for the targets and OARs per our institutional protocols. The mean calculation time for 13 influence matrices in robust optimization by VPMC was 41.6 seconds. Conclusion: VPMC has been successfully enhanced to model aperture blocks in dose calculation and optimization for the PBSPT-based SRS.

physics.med-ph↗

Beam mask and sliding window-facilitated deep learning-based accurate and efficient dose prediction for pencil beam scanning proton therapy

Purpose: To develop a DL-based PBSPT dose prediction workflow with high accuracy and balanced complexity to support on-line adaptive proton therapy clinical decision and subsequent replanning. Methods: PBSPT plans of 103 prostate cancer patients and 83 lung cancer patients previously treated at our institution were included in the study, each with CTs, structure sets, and plan doses calculated by the in-house developed Monte-Carlo dose engine. For the ablation study, we designed three experiments corresponding to the following three methods: 1) Experiment 1, the conventional region of interest (ROI) method. 2) Experiment 2, the beam mask (generated by raytracing of proton beams) method to improve proton dose prediction. 3) Experiment 3, the sliding window method for the model to focus on local details to further improve proton dose prediction. A fully connected 3D-Unet was adopted as the backbone. Dose volume histogram (DVH) indices, 3D Gamma passing rates, and dice coefficients for the structures enclosed by the iso-dose lines between the predicted and the ground truth doses were used as the evaluation metrics. The calculation time for each proton dose prediction was recorded to evaluate the method's efficiency. Results: Compared to the conventional ROI method, the beam mask method improved the agreement of DVH indices for both targets and OARs and the sliding window method further improved the agreement of the DVH indices. For the 3D Gamma passing rates in the target, OARs, and BODY (outside target and OARs), the beam mask method can improve the passing rates in these regions and the sliding window method further improved them. A similar trend was also observed for the dice coefficients. In fact, this trend was especially remarkable for relatively low prescription isodose lines. The dose predictions for all the testing cases were completed within 0.25s.

physics.med-ph↗

Deep-Learning-based Fast and Accurate 3D CT Deformable Image Registration in Lung Cancer

Purpose: In some proton therapy facilities, patient alignment relies on two 2D orthogonal kV images, taken at fixed, oblique angles, as no 3D on-the-bed imaging is available. The visibility of the tumor in kV images is limited since the patient's 3D anatomy is projected onto a 2D plane, especially when the tumor is behind high-density structures such as bones. This can lead to large patient setup errors. A solution is to reconstruct the 3D CT image from the kV images obtained at the treatment isocenter in the treatment position. Methods: An asymmetric autoencoder-like network built with vision-transformer blocks was developed. The data was collected from 1 head and neck patient: 2 orthogonal kV images (1024x1024 voxels), 1 3D CT with padding (512x512x512) acquired from the in-room CT-on-rails before kVs were taken and 2 digitally-reconstructed-radiograph (DRR) images (512x512) based on the CT. We resampled kV images every 8 voxels and DRR and CT every 4 voxels, thus formed a dataset consisting of 262,144 samples, in which the images have a dimension of 128 for each direction. In training, both kV and DRR images were utilized, and the encoder was encouraged to learn the jointed feature map from both kV and DRR images. In testing, only independent kV images were used. The full-size synthetic CT (sCT) was achieved by concatenating the sCTs generated by the model according to their spatial information. The image quality of the synthetic CT (sCT) was evaluated using mean absolute error (MAE) and per-voxel-absolute-CT-number-difference volume histogram (CDVH). Results: The model achieved a speed of 2.1s and a MAE of <40HU. The CDVH showed that <5% of the voxels had a per-voxel-absolute-CT-number-difference larger than 185 HU. Conclusion: A patient-specific vision-transformer-based network was developed and shown to be accurate and efficient to reconstruct 3D CT images from kV images.

cs.CV↗