SearcharxivSearch

arXiv subjects

Hewei Gao

Publications and source records attributed to Hewei Gao.

14 recordsLinked to original sources

Foveated-Imaging Geometry CT Architecture and Seeded Diffusion Model Enabling Global Super-Resolution Reconstruction

For X-ray computed tomography (CT), a smaller detector pixel size generally leads to higher scanner spatial resolution, but inevitably increases system cost as well as data overhead in acquisition and processing. To achieve high-resolution (HR) CT imaging in a more resource-efficient manner, we propose a Foveated-Imaging Geometry CT (FIGCT) architecture, which integrates local HR data into an acquisition scheme dominated by low-resolution (LR) measurements. We further develop a Diffusion Probabilistic FIGCT Super-Resolution Reconstruction (DPFSR) framework to generate global HR CT images over the full field of view (FOV). The concept of FIGCT is first established, and its typical configurations are characterized according to the arrangement of HR data. Two key indices, namely the HR data fraction (HDF) and the LR-to-HR detector pixel size ratio (LHR), are introduced to describe the FIGCT geometry. The proposed DPFSR incorporates local HR information into intermediate clean-image estimates in both the projection and image domains during the reverse diffusion process. This additional step not only guides HR image generation from LR data, but also improves data consistency between the clean-image estimates and the originally measured data. Preliminary numerical simulation results on FIGCT show that the proposed architecture provides high-precision CT images within the region of interest (ROI) corresponding to the HR data, while the spatial resolution deteriorates rapidly outside the ROI. With DPFSR, global HR reconstruction is achieved on the AAPM Grand Challenge dataset and swine lung CT data, outperforming existing SR methods in terms of Learned Perceptual Image Patch Similarity (LPIPS), PSNR, and SSIM.

physics.med-ph

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive ecosystem for music reward modeling under Compositional Multimodal Instruction (CMI), where the generated music may be conditioned on text descriptions, lyrics, and audio prompts. We first introduce CMI-Pref-Pseudo, a large-scale preference dataset comprising 110k pseudo-labeled samples, and CMI-Pref, a high-quality, human-annotated corpus tailored for fine-grained alignment tasks. To unify the evaluation landscape, we propose CMI-RewardBench, a unified benchmark that evaluates music reward models on heterogeneous samples across musicality, text-music alignment, and compositional instruction alignment. Leveraging these resources, we develop CMI reward models (CMI-RMs), a parameter-efficient reward model family capable of processing heterogeneous inputs. We evaluate their correlation with human judgment scores on musicality and alignment on CMI-Pref along with previous datasets. Further experiments demonstrate that CMI-RM not only correlates strongly with human judgments, but also enables effective inference-time scaling via top-k filtering. Code is available at GitHub (https://github.com/Haiwen-Xia/CMI-RewardBench). Model weights: CMI-RM (https://huggingface.co/HaiwenXia/CMI-RM). Datasets: CMI-Pref-Pseudo (https://huggingface.co/datasets/HaiwenXia/cmi-pref-pseudo) and CMI-Pref (https://huggingface.co/datasets/HaiwenXia/cmi-pref)

cs.SD

Matrixed-Spectrum Decomposition Accelerated Linear Boltzmann Transport Equation Solver for Fast Scatter Correction in Multi-Spectral CT

X-ray scatter has been a serious concern in computed tomography (CT), leading to image artifacts and distortion of CT values. The linear Boltzmann transport equation (LBTE) is recognized as a fast and accurate approach for scatter estimation. However, for multi-spectral CT, it is cumbersome to compute multiple scattering components for different spectra separately when applying LBTE-based scatter correction. In this work, we propose a Matrixed-Spectrum Decomposition accelerated LBTE solver (MSD-LBTE) that can be used to compute X-ray scatter distributions from CT acquisitions at two or more different spectra simultaneously, in a unified framework with no sacrifice in accuracy and nearly no increase in computation in theory. First, a matrixed-spectrum solver of LBTE is obtained by introducing an additional label dimension to expand the phase space. Then, we propose a ``spectrum basis'' for LBTE and a principle of selection of basis using the QR decomposition, along with the above solver to construct the MSD-LBTE. Based on MSD-LBTE, a unified scatter correction method can be established for multi-spectral CT. We validate the effectiveness and accuracy of our method by comparing it with the Monte Carlo method, including the computational time. We also evaluate the scatter correction performance using two different phantoms for fast-kV switching based dual-energy CT, and using an elliptical phantom in a numerical simulation for kV-modulation enabled CT scans, validating that our proposed method can significantly reduce the computational cost at multiple spectra and effectively reduce scatter artifact in reconstructed CT images.

physics.med-ph

FedNano: Toward Lightweight Federated Tuning for Pretrained Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) excel in tasks like multimodal reasoning and cross-modal retrieval but face deployment challenges in real-world scenarios due to distributed multimodal data and strict privacy requirements. Federated Learning (FL) offers a solution by enabling collaborative model training without centralizing data. However, realizing FL for MLLMs presents significant challenges, including high computational demands, limited client capacity, substantial communication costs, and heterogeneous client data. Existing FL methods assume client-side deployment of full models, an assumption that breaks down for large-scale MLLMs due to their massive size and communication demands. To address these limitations, we propose FedNano, the first FL framework that centralizes the LLM on the server while introducing NanoEdge, a lightweight module for client-specific adaptation. NanoEdge employs modality-specific encoders, connectors, and trainable NanoAdapters with low-rank adaptation. This design eliminates the need to deploy LLM on clients, reducing client-side storage by 95%, and limiting communication overhead to only 0.01% of the model parameters. By transmitting only compact NanoAdapter updates, FedNano handles heterogeneous client data and resource constraints while preserving privacy. Experiments demonstrate that FedNano outperforms prior FL baselines, bridging the gap between MLLM scale and FL feasibility, and enabling scalable, decentralized multimodal AI systems.

cs.LG

Energy-Threshold Bias Calculator: A Physics-Model Based Adaptive Correction Scheme for Photon-Counting CT

Photon-counting detector based computed tomography (PCCT) has greatly advanced in recent years. However, spectral inconsistency, referring to inter-pixel variations in detected counts per energy bin, can easily leads to ring or band artifacts and inaccuracies in CT reconstructed images. This work proposes a novel physics-model based method to correct for spectral inconsistency by modeling it through two terms: (1) a fixed spectral skew term (energy threshold-independent filtration function) determined at a given energy threshold, and (2) a variable energy-threshold bias term that can be directly calculated by using our spectral model as the threshold changes. After the two terms being computed out in the calibration stage, they will be incorporated into our spectral model to adaptively generate the spectral correction vectors as well as the material decomposition vectors if needed, pixel-by-pixel for PCCT projection data. Using a minimum set of parameters with explicit physics meaning, such an energy-threshold bias calculator (ETB-Cal) has advantages of computational efficiency, robustness in implementation, and convenience with no need of X-ray fluorescence materials in calibration. To validate our method, both numerical simulations and physical experiments using multiple phantoms were carried out on a tabletop PCCT system, with preliminary results showing a significant reduction in non-uniformity, from 29.3 to 5.8 HU for Gammex multi-energy phantom versus no correction (comparatively, 8.3 HU was achieved by a polynomial-involving model-based approach with no explicit modeling and calculating of energy threshold bias but more calibration data required), and from 27.9 to 3.2 HU for the Kyoto head phantom.

physics.med-ph

ComptoNet: An End-to-End Deep Learning Framework for Scatter Estimation in Multi-Source Stationary CT

Multi-source stationary computed tomography (MSS-CT) offers significant advantages in medical and industrial applications due to its gantry-less scan architecture and/or capability of simultaneous multi-source emission. However, the lack of anti-scatter grid deployment in MSS-CT results in severe forward and/or cross scatter contamination, presenting a critical challenge that necessitates an accurate and efficient scatter correction. In this work, ComptoNet, an innovative end-to-end deep learning framework for scatter estimation in MSS-CT, is proposed, which integrates Compton-scattering physics with deep learning techniques to address the challenges of scatter estimation effectively. Central to ComptoNet is the Compton-map, a novel concept that captures the distribution of scatter signals outside the scan field of view, primarily consisting of large-angle Compton scatter. In ComptoNet, a reference Compton-map and/or spare detector data are used to guide the physics-driven deep estimation of scatter from simultaneous emissions by multiple sources. Additionally, a frequency attention module is employed for enhancing the low-frequency smoothness. Such a multi-source deep scatter estimation framework decouples the cross and forward scatter. It reduces network complexity and ensures a consistent low-frequency signature with different photon numbers of simulations, as evidenced by mean absolute percentage errors (MAPEs) that are less than $1.26\%$. Conducted by using data generated from Monte Carlo simulations with various phantoms, experiments demonstrate the effectiveness of ComptoNet, with significant improvements in scatter estimation accuracy (a MAPE of $0.84\%$). After scatter correction, nearly artifact-free CT images are obtained, further validating the capability of our proposed ComptoNet in mitigating scatter-induced errors.

physics.med-ph

Penumbra-Effect Induced Spectral Mixing in X-ray Computed Tomography: A Multi-Ray Spectrum Estimation Model and Subsampled Weighting Algorithm

Purpose: With the development of spectral CT, several novel spectral filters have been introduced to modulate the spectra, such as split filters and spectral modulators. However, due to the finite size of the focal spot of X-ray source, these filters cause spectral mixing in the penumbra region. Traditional spectrum estimation methods fail to account for it, resulting in reduced spectral accuracy. Methods: To address this challenge, we develop a multi-ray spectrum estimation model and propose an Adaptive Subsampled WeIghting of Filter Thickness (A-SWIFT) method. First, we estimate the unfiltered spectrum using traditional methods. Next, we model the final spectra as a weighted summation of spectra attenuated by multiple filters. The weights and equivalent lengths are obtained by X-ray transmission measurements taken with altered spectra using different kVp or flat filters. Finally, the spectra are approximated by using the multi-ray model. To mimic the penumbra effect, we used a spectral modulator (0.2 mm Mo, 0.6 mm Mo) and a split filter (0.07 mm Au, 0.7 mm Sn) in simulations, and used a copper modulator and a molybdenum modulator (0.2 mm, 0.6 mm) in experiments. Results: Simulation results show that the mean energy bias in the penumbra region decreased from 7.43 keV using the previous SCFM method (Spectral Compensation for Modulator) to 0.72 keV using the A-SWIFT method for the split filter, and from 1.98 keV to 0.61 keV for the spectral modulator. In experiments, the root mean square error of the selected ROIs was decreased from 77 to 7 Hounsfield units (HU) for the pure water phantom with a molybdenum modulator, and from 85 to 21 HU with a copper modulator. Conclusion: Based on a multi-ray spectrum estimation model, the A-SWIFT method provides an accurate and robust approach for spectrum estimation in penumbra region of CT systems utilizing spectral filters.

physics.med-ph

Fast KV-Switching and Dual-Layer Flat-Panel Detector Enabled Cone-Beam CT Joint Spectral Imaging

Purpose: Fast kV-switching (FKS) and dual-layer flat-panel detector (DL-FPD) technologies have been actively studied as promising dual-energy solutions for FPD-based cone-beam computed tomography (CBCT). However, CBCT spectral imaging is known to face challenges in obtaining accurate and robust material discrimination performance due to the limited energy separation. To further improve CBCT spectral imaging capability, this work aims to promote a source-detector joint spectral imaging solution which takes advantages of both FKS and DL-FPD, and to conduct a feasibility study on the first tabletop CBCT system with the joint spectral imaging capability developed. Methods: In this work, the first FKS and DL-FPD jointly enabled multi-energy tabletop CBCT system has been developed in our laboratory. To evaluate its spectral imaging performance, a set of physics experiments are conducted, where the multi-energy and head phantoms are scanned using the 80/105/130kVp switching pairs and projection data are collected using a prototype DL-FPD. To compensate for the slightly angular mismatch between the low- and high-energy projections in FKS, a dual-domain projection completion scheme is implemented. Afterwards material decomposition is carried out by using the maximum-likelihood method, followed by reconstruction of basis material and virtual monochromatic images. Results: The physics experiments confirmed the feasibility and superiority of the joint spectral imaging, whose CNR of the multi-energy phantom were boosted by an average improvement of 21.9%, 20.4% for water and 32.8%, 62.8% for iodine when compared with that of the FKS and DL-FPD in fan-beam and cone-beam experiments, respectively. Conclusions: A feasibility study of the joint spectral imaging for CBCT by utilizing both the FKS and DL-FPD was conducted, with the first tabletop CBCT system having such a capability being developed.

physics.med-ph

Generalized-Equiangular Geometry CT: Concept and Shift-Invariant FBP Algorithms

With advanced X-ray source and detector technologies being continuously developed, non-traditional CT geometries have been widely explored. Generalized-Equiangular Geometry CT (GEGCT) architecture, in which an X-ray source might be positioned radially far away from the focus of arced detector array that is equiangularly spaced, is of importance in many novel CT systems and designs. GEGCT, unfortunately, has no theoretically exact and shift-invariant analytical image reconstruction algorithm in general. In this study, to obtain fast and accurate reconstruction from GEGCT and to promote its system design and optimization, an in-depth investigation on a group of approximate Filtered BackProjection (FBP) algorithms with a variety of weighting strategies has been conducted. The architecture of GEGCT is first presented and characterized by using a normalized-radial-offset distance (NROD). Next, shift-invariant weighted FBP-type algorithms are derived in a unified framework, with pre-filtering, filtering, and post-filtering weights. Three viable weighting strategies are then presented including a classic one developed by Besson in the literature and two new ones generated from a curvature fitting and from an empirical formula, where all of the three weights can be expressed as certain functions of NROD. After that, an analysis of reconstruction accuracy is conducted with a wide range of NROD. We further stretch the weighted FBP-type algorithms to GEGCT with dynamic NROD. Finally, the weighted FBP algorithm for GEGCT is extended to a three-dimensional form in the case of cone-beam scan with a cylindrical detector array.

physics.med-ph

Multi-Energy Blended CBCT Spectral Imaging Using a Spectral Modulator with Flying Focal Spot (SMFFS)

Conventional cone-beam CT (CBCT) can be easily compromised by scatter and beam hardening artifacts, and the entanglement of scatter and spectral effects introduces additional complexity. In this work, we present the first attempt to develop a stationary spectral modulator with flying focal spot (SMFFS) technology as a promising, low-cost approach to accurately solving the X-ray scattering problem and physically enabling spectral imaging in a unified framework. To deal with the intertwined scatter-spectral challenge, we propose a novel scatter-decoupled material decomposition (SDMD) method for SMFFS based on a hypothesis of scatter similarity. Monte Carlo simulations of a pure-water cylinder phantom with different focal spot deflections show that focal spot deflections within a range of ~2 mm share quite similar scatter distributions overall. Numerical simulations using a clinical abdominal CT dataset demonstrate that SMFFS with SDMD method can achieve better material decomposition and CT number accuracy with less artifacts. Physics experiments on a tabletop CBCT system using a Gammex multi-energy CT phantom an anthropomorphic chest phantom, are carried out to demonstrate the feasibility of CBCT spectral imaging with SMFFS. For the chest phantom, the root mean square error (RMSE) in selected regions of interest (ROIs) of virtual monochromatic image (VMI) at 70 keV is 11.8 HU for SMFFS CB scan, and 14.5 and 437.6 HU for sequential 80/140 kVp (DKV) CB scan with and without scatter correction, respectively. Also, the non-uniformity among selected regions is 14.1 HU for SMFFS CB scan, and 59.4 and 184.0 HU for the DKV CB scan with and without a traditional scatter correction method, respectively. Our preliminary results show that SMFFS can enable spectral imaging with simultaneous scatter correction for CBCT and effectively improve its quantitative imaging performance.

physics.med-ph

Fluence Adaptation for Task-based Dose Optimization in X-ray Phase-Contrast Imaging

Purpose: Grating-based imaging (GBI) and edge-illumination (EI) are two promising types of XPCI as the conventional x-ray sources can be directly utilized. For GBI and EI systems, the phase-stepping acquisition with multiple exposures at a constant fluence is usually adopted in the literature. This work, however, attempts to challenge such a constant fluence concept during the phase-stepping process and proposes a fluence adaptation mechanism for dose reduction. Method: Recently, analytic multi-order moment analysis has been proposed to improve the computing efficiency. In these algorithms, multiple contrasts can be calculated by summing together the weighted phase-stepping curves (PSCs) with some kernel functions, which suggests us that the raw data at different steps have different contributions for the noise in retrieved contrasts. Based on analytic retrieval formulas and the Gaussian noise model for detected signals, we derived an optimal adaptive fluence distribution, which is proportional to the absolute weighting kernel functions and the root of original sample PSCs acquired under the constant fluence. Results: To validate our analyses, simulations and experiments are conducted for GBI and EI systems. Simulated results demonstrate that the dose reduction ratio between our proposed fluence distributions and the typical constant one can be about 20% for the phase contrast, which is consistent with our theoretical predictions. Although the experimental noise reduction ratios are a little smaller than the theoretical ones, synthetic and real experiments both observe better noise performance by our proposed method. Our simulated results also give out the effective ranges of the parameters of the PSCs, such as the visibility in GBI, the standard deviation and the mean value in EI, providing a guidance for the use of our proposed approach in practice.

physics.med-ph

An Analysis of Scatter Characteristics in X-ray CT Spectral Correction

X-ray scatter remains a major physics challenge in volumetric computed tomography (CT), whose physical and statistical behaviors have been commonly leveraged in order to eliminate its impact on CT image quality. In this work, we conduct an in-depth derivation of how the scatter distribution and scatter to primary ratio (SPR) will change during the spectral correction, leading to an interesting finding on the property of scatter: when applying the spectral correction before scatter is removed, the impact of SPR on a CT projection will be scaled by the first derivative of the mapping function; while the scatter distribution in the transmission domain will be scaled by the product of the first derivative of the mapping function and a natural exponential of the projection difference before and after the mapping. Such a characterization of scatter's behavior provides an analytic approach of compensating for the SPR as well as approximating the change of scatter distribution after spectral correction, even though both of them might be significantly distorted as the linearization mapping function in spectral correction could vary a lot from one detector pixel to another. We conduct an evaluation of SPR compensations on a Catphan phantom and an anthropomorphic chest phantom to validate the characteristics of scatter. In addition, this scatter property is also directly adopted into CT imaging using a spectral modulator with flying focal spot technology (SMFFS) as an example to demonstrate its potential in practical applications.

physics.med-ph

Truncated analytic moment analysis and its hybrid-field contrast in grating-based x-ray phase contrast imaging

For grating-based x-ray phase contrast imaging (GPCI), a multi-order moment analysis (MMA) has been recently developed to obtain multiple contrasts from the ultra-small-angle x-ray scattering distribution, as a novel information retrieval approach that is totally different from the conventional Fourier components analysis (FCA). In this paper, we present an analytic form of MMA in theory that can retrieve multiple contrasts directly from raw phase-stepping images, with no scattering distribution involved. For practical implementation, a truncated analytic analysis (called as TA-MMA) is adopted and it is hundreds of times faster in computation than the original deconvolution-based MMA (called as DB-MMA). More importantly, TA-MMA is proved to establish a quantitative connection between FCA and MMA, i.e., the first-order moment computed by TA-MMA is essentially the product of the phase contrast and the dark-field contrast retrieved by FCA, providing a new physical parameter for GPCI. The new physical parameter, in fact, can be treated as a ``hybrid-field'' contrast as it fuses the original phase contrast and dark-field contrast in a straightforward manner, which may be the first physical fusion contrast in GPCI to our knowledge and may have a potential to be directly used in practical applications.

physics.med-ph

A cascaded dual-domain deep learning reconstruction method for sparsely spaced multidetector helical CT

Helical CT has been widely used in clinical diagnosis. Sparsely spaced multidetector in z direction can increase the coverage of the detector provided limited detector rows. It can speed up volumetric CT scan, lower the radiation dose and reduce motion artifacts. However, it leads to insufficient data for reconstruction. That means reconstructions from general analytical methods will have severe artifacts. Iterative reconstruction methods might be able to deal with this situation but with the cost of huge computational load. In this work, we propose a cascaded dual-domain deep learning method that completes both data transformation in projection domain and error reduction in image domain. First, a convolutional neural network (CNN) in projection domain is constructed to estimate missing helical projection data and converting helical projection data to 2D fan-beam projection data. This step is to suppress helical artifacts and reduce the following computational cost. Then, an analytical linear operator is followed to transfer the data from projection domain to image domain. Finally, an image domain CNN is added to improve image quality further. These three steps work as an entirety and can be trained end to end. The overall network is trained using a simulated lung CT dataset with Poisson noise from 25 patients. We evaluate the trained network on another three patients and obtain very encouraging results with both visual examination and quantitative comparison. The resulting RRMSE is 6.56% and the SSIM is 99.60%. In addition, we test the trained network on the lung CT dataset with different noise level and a new dental CT dataset to demonstrate the generalization and robustness of our method.

physics.med-ph