SearcharxivSearch

arXiv subjects

Yansong Zhu

Publications and source records attributed to Yansong Zhu.

6 recordsLinked to original sources

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression

Interactive video world models commonly convert bidirectional video generators into causal autoregressive systems through control fine-tuning, autoregressive training, causal initialization, and few-step distillation. This pipeline is costly, while frozen causal histories accumulate errors that degrade long-horizon fidelity and controllability. We present BiWM, the first open-source full-stack training framework for bidirectional autoregressive video world models. BiWM retains full attention within each generated chunk and requires only two stages: camera/action-control fine-tuning and few-step Distribution Matching Distillation (DMD). Both stages converge within a few hundred optimizer steps on 8 H200 GPUs. The framework supports Wan2.1-T2V-1.3B, Wan2.2-TI2V-5B, HunyuanVideo-1.5-TI2V-8B, and LTX-2.3-22B, together with real-world camera control, pluggable long-history compression, and optional low-bit deployment. Supervised and forward-KL anchors mitigate DMD mode collapse and preserve scene dynamics. BiWM provides a compact, reproducible path from pretrained bidirectional video models to interactive, controllable, and efficient world models.

cs.CV

Feasibility of PET-enabled dual-energy CT imaging: First physical phantom and initial patient results

X-ray computed tomography (CT) in PET/CT is commonly operated with a single energy, resulting in a limitation of lacking tissue composition information. Dual-energy (DE) spectral CT enables material decomposition by using two different x-ray energies and may be combined with PET for improved multimodality imaging, but would either require hardware upgrade or increase radiation dose due to the added second x-ray CT scan. Recently proposed PET-enabled DECT method allows dual-energy spectral imaging using a conventional PET/CT scanner without the need for a second x-ray CT scan. A gamma-ray CT (gCT) image at 511 keV can be generated from the existing time-of-flight PET data with the maximum-likelihood attenuation and activity (MLAA) approach and is then combined with the low-energy x-ray CT image to form dual-energy spectral imaging. To improve the image quality of gCT, a kernel MLAA method was further proposed by incorporating x-ray CT as a priori information. The concept of this PET-enabled DECT has been validated using simulation studies, but not yet with 3D real data. In this work, we developed a general open-source implementation for gCT reconstruction from PET data and use this implementation for the first real data validation with both a physical phantom study and a human subject study on a uEXPLORER total-body PET/CT system. These results have demonstrated the feasibility of this method for spectral imaging and material decomposition.

physics.med-ph

Single-Subject Deep-Learning Image Reconstruction with a Neural Optimization Transfer Algorithm for PET-enabled Dual-Energy CT Imaging

Combining dual-energy computed tomography (DECT) with positron emission tomography (PET) offers many potential clinical applications but typically requires expensive hardware upgrades or increases radiation doses on PET/CT scanners due to an extra X-ray CT scan. The recent PET-enabled DECT method allows DECT imaging on PET/CT without requiring a second X-ray CT scan. It combines the already existing X-ray CT image with a 511 keV γ-ray CT (gCT) image reconstructed from time-of-flight PET emission data. A kernelized framework has been developed for reconstructing gCT image but this method has not fully exploited the potential of prior knowledge. Use of deep neural networks may explore the power of deep learning in this application. However, common approaches require a large database for training, which is impractical for a new imaging method like PET-enabled DECT. Here, we propose a single-subject method by using neural-network representation as a deep coefficient prior to improving gCT image reconstruction without population-based pre-training. The resulting optimization problem becomes the tomographic estimation of nonlinear neural-network parameters from gCT projection data. This complicated problem can be efficiently solved by utilizing the optimization transfer strategy with quadratic surrogates. Each iteration of the proposed neural optimization transfer algorithm includes: PET activity image update; gCT image update; and least-square neural-network learning in the gCT image domain. This algorithm is guaranteed to monotonically increase the data likelihood. Results from computer simulation, real phantom data and real patient data have demonstrated that the proposed method can significantly improve gCT image quality and consequent multi-material decomposition as compared to other methods.

physics.med-ph

Fisher information analysis of list-mode SPECT emission data for joint estimation of activity and attenuation distribution

The potential to perform attenuation and scatter compensation (ASC) in single-photon emission computed tomography (SPECT) imaging using only the SPECT emission data is highly significant. In this context, attenuation in SPECT is primarily due to Compton scattering, where the probability of Compton scatter is proportional to the attenuation coefficient of the tissue and the energy of the scattered photon and the scattering angle are related. Given this premise, we investigate whether the SPECT scattered-photon data acquired in list-mode (LM) format and including the energy information can be used to estimate the attenuation map. For this purpose, we propose a Fisher-information-based method that yields the Cramer-Rao bound (CRB) for the task of jointly estimating the activity/attenuation distribution using only the SPECT emission data. The proposed method is applied to analyze the information content of SPECT LM emission data in a 2D SPECT system using computational studies with digital phantoms for different photon-count levels. The results show that scattered photons contain information to estimate the attenuation coefficients. An increase in the number of detected photons leads to lower CRB for both the attenuation and activity coefficients. Also, the CRB obtained for the attenuation and activity coefficients is typically much lower than the true value of these coefficients. Further, processing the emission data in LM format yields a lower CRB in comparison to binning data. Finally, we observe that systems with better energy resolution yield a lower CRB for the attenuation coefficient. Overall, the results provide strong evidence that LM SPECT emission data, including the scattered photons, contains information to jointly estimate the activity and attenuation coefficients.

physics.med-ph

Image reconstruction in fluorescence molecular tomography with sparsity-initialized maximum-likelihood expectation maximization

We present a reconstruction method involving maximum-likelihood expectation maximization (MLEM) to model Poisson noise as applied to fluorescence molecular tomography (FMT). MLEM is initialized with the output from a sparse reconstruction-based approach, which performs truncated singular value decomposition-based preconditioning followed by fast iterative shrinkage-thresholding algorithm (FISTA) to enforce sparsity. The motivation for this approach is that sparsity information could be accounted for within the initialization, while MLEM would accurately model Poisson noise in the FMT system. Simulation experiments show the proposed method significantly improves images qualitatively and quantitatively. The method results in over 20 times faster convergence compared to uniformly initialized MLEM and improves robustness to noise compared to pure sparse reconstruction. We also theoretically justify the ability of the proposed approach to reduce noise in the background region compared to pure sparse reconstruction. Overall, these results provide strong evidence to model Poisson noise in FMT reconstruction and for application of the proposed reconstruction framework to FMT imaging.

physics.med-ph

Application of computational breast phantoms to evaluate reconstruction methods for fluorescence molecular tomography

Fluorescence molecular tomography (FMT) has potential of providing high contrast images for breast tumor detection. Computational phantom provides a convenient way to a wide variety of fluorophore distribution configurations in patients and perform comprehensive evaluation of the imaging systems and methods for FMT. In this study, a digital breast phantom was used to compare the performance of a novel sparsity-based reconstruction method and Tikhonov regularization method for resolving tumors with different amount of separation. The results showed that the proposed sparse reconstruction method yielded better performance. This simulation-based approach with computational phantoms enabled an evaluation of the reconstruction methods for FMT for breast-cancer detection.

physics.med-ph