Searcharxiv⌕ Search

arXiv subjects

Zhentao Liu

Publications and source records attributed to Zhentao Liu.

At least 19 recordsLinked to original sources

3D Vessel Reconstruction from Sparse-View Dynamic DSA Images via Vessel Probability Guided Attenuation Learning

Digital Subtraction Angiography (DSA) is one of the gold standards for vascular disease diagnosis. With the help of a contrast agent, time-resolved 2D DSA images deliver comprehensive blood flow information and can be utilized to reconstruct 3D vessel structures for medical assessment. Current commercial DSA systems typically require hundreds of scanning views to perform reconstruction, resulting in substantial radiation exposure. In this study, we propose a neural rendering-based optimization framework tailored for high-quality sparse-view DSA reconstruction to reduce radiation dosage. Our approach, termed vessel probability guided attenuation learning, represents DSA imaging as a complementary weighted combination of static and dynamic attenuation fields, with the weights derived from the time-independent vessel probability field. Functioning as a foreground mask, vessel probability provides proper gradients for both static and dynamic fields adaptive to different scene types. This mechanism enables self-supervised decomposition between static backgrounds and dynamic contrast agent flow, and significantly improves reconstruction quality. Our model is trained by minimizing the discrepancy between synthesized projections and real captured DSA images. We further employ two training strategies to improve reconstruction quality: (1) coarse-to-fine progressive training for better geometry and (2) temporal perturbed rendering loss for temporal consistency. Experimental results have demonstrated high-quality 3D vessel reconstruction and 2D DSA image synthesis.

eess.IV↗

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection

The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive detection methods, we propose StreamMark, a novel deep learning-based, semi-fragile audio watermarking system. StreamMark is designed to be robust against benign audio conversions that preserve semantic meaning (e.g., compression, noise) while remaining fragile to malicious, semantics-altering manipulations (e.g., voice conversion, speech editing). Our method introduces a complex-domain embedding technique within a unique Encoder-Distortion-Decoder architecture, trained explicitly to differentiate between these two classes of transformations. Comprehensive benchmarks demonstrate that StreamMark achieves high imperceptibility (SNR 24.16 dB, PESQ 4.20), is resilient to real-world distortions like Opus encoding, and exhibits principled fragility against a suite of deepfake attacks, with message recovery accuracy dropping to chance levels (~50%), while remaining robust to benign AI-based style transfers (ACC >98%).

eess.AS↗

DSA-SRGS: Super-Resolution Gaussian Splatting for Dynamic Sparse-View DSA Reconstruction

Digital subtraction angiography (DSA) is a key imaging technique for the auxiliary diagnosis and treatment of cerebrovascular diseases. Recent advancements in gaussian splatting and dynamic neural representations have enabled robust 3D vessel reconstruction from sparse dynamic inputs. However, these methods are fundamentally constrained by the resolution of input projections, where performing naive upsampling to enhance rendering resolution inevitably results in severe blurring and aliasing artifacts. Such lack of super-resolution capability prevents the reconstructed 4D models from recovering fine-grained vascular details and intricate branching structures, which restricts their application in precision diagnosis and treatment. To solve this problem, this paper proposes DSA-SRGS, the first super-resolution gaussian splatting framework for dynamic sparse-view DSA reconstruction. Specifically, we introduce a Multi-Fidelity Texture Learning Module that integrates high-quality priors from a fine-tuned DSA-specific super-resolution model, into the 4D reconstruction optimization. To mitigate potential hallucination artifacts from pseudo-labels, this module employs a Confidence-Aware Strategy to adaptively weight supervision signals between the original low-resolution projections and the generated high-resolution pseudo-labels. Furthermore, we develop Radiative Sub-Pixel Densification, an adaptive strategy that leverages gradient accumulation from high-resolution sub-pixel sampling to refine the 4D radiative gaussian kernels. Extensive experiments on two clinical DSA datasets demonstrate that DSA-SRGS significantly outperforms state-of-the-art methods in both quantitative metrics and qualitative visual fidelity.

cs.CV↗

Adaptive information-maximization encoding for ghost imaging--A general Bayesian framework under experimental physical constraints

Ghost imaging (GI) has demonstrated diverse imaging capabilities enabled by its encoding-decoding-based computational imaging mechanism. Accordingly, information-theoretic studies have emerged as a promising avenue for probing the fundamental performance bounds of of GI and related computational imaging paradigms. However, the design of information-theoretically optimal encoding strategies remains largely unexplored, primarily due to the intractability of the prior probability density function (PDF) of an unknown scene. Here, by leveraging the ability of recursively estimating the PDF of the object to be imaged via Bayesian filtering, we propose to establish an adaptive information-maximization encoding (AIME) design framework. Based on the adaptively estimated posterior PDF from previously acquired measurements, the expected information gain of subsequent detections is evaluated and maximized to design the corresponding encoding patterns in a closed-loop manner. Within this framework, the theoretical form of the information-optimal encoding under representative physical constraints is analytically derived. Corresponding experimental results show that, GI systems employing information-optimal encoding achieve markedly improved imaging performance compared with conventional fixed point-to-point imaging without relying on additional heuristic regularization schemes, particularly in low signal-to-noise ratio regimes. Moreover, the proposed strategy consistently enables significantly enhanced information acquisition capability compared with existing encoding strategies, leading to substantially improved imaging quality. These results establish a principled information-theoretic foundation for optimal encoding design in computational imaging paradigms,provided that the forward model can be accurately characterized.

physics.optics↗

3D MedDiffusion: A 3D Medical Latent Diffusion Model for Controllable and High-quality Medical Image Generation

The generation of medical images presents significant challenges due to their high-resolution and three-dimensional nature. Existing methods often yield suboptimal performance in generating high-quality 3D medical images, and there is currently no universal generative framework for medical imaging. In this paper, we introduce a 3D Medical Latent Diffusion (3D MedDiffusion) model for controllable, high-quality 3D medical image generation. 3D MedDiffusion incorporates a novel, highly efficient Patch-Volume Autoencoder that compresses medical images into latent space through patch-wise encoding and recovers back into image space through volume-wise decoding. Additionally, we design a new noise estimator to capture both local details and global structural information during diffusion denoising process. 3D MedDiffusion can generate fine-detailed, high-resolution images (up to 512x512x512) and effectively adapt to various downstream tasks as it is trained on large-scale datasets covering CT and MRI modalities and different anatomical regions (from head to leg). Experimental results demonstrate that 3D MedDiffusion surpasses state-of-the-art methods in generative quality and exhibits strong generalizability across tasks such as sparse-view CT reconstruction, fast MRI reconstruction, and data augmentation for segmentation and classification. Source code and checkpoints are available at https://github.com/ShanghaiTech-IMPACT/3D-MedDiffusion.

eess.IV↗

ST2HE: A Cross-Platform Framework for Virtual Histology and Annotation of High-Resolution Spatial Transcriptomics Data

High-resolution spatial transcriptomics (HR-ST) technologies offer unprecedented insights into tissue architecture but lack standardized frameworks for histological annotation. We present ST2HE, a cross-platform generative framework that synthesizes virtual hematoxylin and eosin (H&E) images directly from HR-ST data. ST2HE integrates nuclei morphology and spatial transcript coordinates using a one-step diffusion model, enabling histologically faithful image generation across diverse tissue types and HR-ST platforms. Conditional and tissue-independent variants support both known and novel tissue contexts. Evaluations on breast cancer, non-small cell lung cancer, and Kaposi's sarcoma demonstrate ST2HE's ability to preserve morphological features and support downstream annotations of tissue histology and phenotype classification. Ablation studies reveal that larger context windows, balanced loss functions, and multi-colored transcript visualization enhance image fidelity. ST2HE bridges molecular and histological domains, enabling interpretable, scalable annotation of HR-ST data and advancing computational pathology.

q-bio.QM↗

4DRGS: 4D Radiative Gaussian Splatting for Efficient 3D Vessel Reconstruction from Sparse-View Dynamic DSA Images

Reconstructing 3D vessel structures from sparse-view dynamic digital subtraction angiography (DSA) images enables accurate medical assessment while reducing radiation exposure. Existing methods often produce suboptimal results or require excessive computation time. In this work, we propose 4D radiative Gaussian splatting (4DRGS) to achieve high-quality reconstruction efficiently. In detail, we represent the vessels with 4D radiative Gaussian kernels. Each kernel has time-invariant geometry parameters, including position, rotation, and scale, to model static vessel structures. The time-dependent central attenuation of each kernel is predicted from a compact neural network to capture the temporal varying response of contrast agent flow. We splat these Gaussian kernels to synthesize DSA images via X-ray rasterization and optimize the model with real captured ones. The final 3D vessel volume is voxelized from the well-trained kernels. Moreover, we introduce accumulated attenuation pruning and bounded scaling activation to improve reconstruction quality. Extensive experiments on real-world patient data demonstrate that 4DRGS achieves impressive results in 5 minutes training, which is 32x faster than the state-of-the-art method. This underscores the potential of 4DRGS for real-world clinics.

eess.IV↗

Hyperspectral image reconstruction by deep learning with super-Rayleigh speckles

Ghost imaging via sparsity constraints (GISC) spectral camera modulates the three-dimensional (3D) hyperspectral image into a two-dimensional (2D) compressive image with speckles in a single shot. It obtains a 3D hyperspectral image (HSI) by reconstruction algorithms. The rapid development of deep learning has provided a new method for 3D HSI reconstruction. Moreover, the imaging performance of the GISC spectral camera can be improved by optimizing the speckle modulation. In this paper, we propose an end-to-end GISCnet with super-Rayleigh speckle modulation to improve the imaging quality of the GISC spectral camera. The structure of GISCnet is very simple but effective, and we can easily adjust the network structure parameters to improve the image reconstruction quality. Relative to Rayleigh speckles, our super-Rayleigh speckles modulation exhibits a wealth of detail in reconstructing 3D HSIs. After evaluating 648 3D HSIs, it was found that the average peak signal-to-noise ratio increased from 27 dB to 31 dB. Overall, the proposed GISCnet with super-Rayleigh speckle modulation can effectively improve the imaging quality of the GISC spectral camera by taking advantage of both optimized super-Rayleigh modulation and deep-learning image reconstruction, inspiring joint optimization of light-field modulation and image reconstruction to improve ghost imaging performance.

eess.IV↗

Geometry-Aware Attenuation Learning for Sparse-View CBCT Reconstruction

Cone Beam Computed Tomography (CBCT) plays a vital role in clinical imaging. Traditional methods typically require hundreds of 2D X-ray projections to reconstruct a high-quality 3D CBCT image, leading to considerable radiation exposure. This has led to a growing interest in sparse-view CBCT reconstruction to reduce radiation doses. While recent advances, including deep learning and neural rendering algorithms, have made strides in this area, these methods either produce unsatisfactory results or suffer from time inefficiency of individual optimization. In this paper, we introduce a novel geometry-aware encoder-decoder framework to solve this problem. Our framework starts by encoding multi-view 2D features from various 2D X-ray projections with a 2D CNN encoder. Leveraging the geometry of CBCT scanning, it then back-projects the multi-view 2D features into the 3D space to formulate a comprehensive volumetric feature map, followed by a 3D CNN decoder to recover 3D CBCT image. Importantly, our approach respects the geometric relationship between 3D CBCT image and its 2D X-ray projections during feature back projection stage, and enjoys the prior knowledge learned from the data population. This ensures its adaptability in dealing with extremly sparse view inputs without individual training, such as scenarios with only 5 or 10 X-ray projections. Extensive evaluations on two simulated datasets and one real-world dataset demonstrate exceptional reconstruction quality and time efficiency of our method.

eess.IV↗

TeethDreamer: 3D Teeth Reconstruction from Five Intra-oral Photographs

Orthodontic treatment usually requires regular face-to-face examinations to monitor dental conditions of the patients. When in-person diagnosis is not feasible, an alternative is to utilize five intra-oral photographs for remote dental monitoring. However, it lacks of 3D information, and how to reconstruct 3D dental models from such sparse view photographs is a challenging problem. In this study, we propose a 3D teeth reconstruction framework, named TeethDreamer, aiming to restore the shape and position of the upper and lower teeth. Given five intra-oral photographs, our approach first leverages a large diffusion model's prior knowledge to generate novel multi-view images with known poses to address sparse inputs and then reconstructs high-quality 3D teeth models by neural surface reconstruction. To ensure the 3D consistency across generated views, we integrate a 3D-aware feature attention mechanism in the reverse diffusion process. Moreover, a geometry-aware normal loss is incorporated into the teeth reconstruction process to enhance geometry accuracy. Extensive experiments demonstrate the superiority of our method over current state-of-the-arts, giving the potential to monitor orthodontic treatment remotely. Our code is available at https://github.com/ShanghaiTech-IMPACT/TeethDreamer

cs.CV↗

Multi-View Vertebra Localization and Identification from CT Images

Accurately localizing and identifying vertebrae from CT images is crucial for various clinical applications. However, most existing efforts are performed on 3D with cropping patch operation, suffering from the large computation costs and limited global information. In this paper, we propose a multi-view vertebra localization and identification from CT images, converting the 3D problem into a 2D localization and identification task on different views. Without the limitation of the 3D cropped patch, our method can learn the multi-view global information naturally. Moreover, to better capture the anatomical structure information from different view perspectives, a multi-view contrastive learning strategy is developed to pre-train the backbone. Additionally, we further propose a Sequence Loss to maintain the sequential structure embedded along the vertebrae. Evaluation results demonstrate that, with only two 2D networks, our method can localize and identify vertebrae in CT images accurately, and outperforms the state-of-the-art methods consistently. Our code is available at https://github.com/ShanghaiTech-IMPACT/Multi-View-Vertebra-Localization-and-Identification-from-CT-Images.

eess.IV↗

Electrically programmable magnetic coupling in an Ising network exploiting solid-state ionic gating

Two-dimensional arrays of magnetically coupled nanomagnets provide a mesoscopic platform for exploring collective phenomena as well as realizing a broad range of spintronic devices. In particular, the magnetic coupling plays a critical role in determining the nature of the cooperative behaviour and providing new functionalities in nanomagnet-based devices. Here, we create coupled Ising-like nanomagnets in which the coupling between adjacent nanomagnetic regions can be reversibly converted between parallel and antiparallel through solid-state ionic gating. This is achieved with the voltage-control of magnetic anisotropies in a nanosized region where the symmetric exchange interaction favours parallel alignment and the antisymmetric exchange interaction, namely the Dzyaloshinskii-Moriya interaction, favours antiparallel alignment. Applying this concept to a two-dimensional lattice, we demonstrate a voltage-controlled phase transition in artificial spin ices. Furthermore, we achieve an addressable control of the individual couplings and realize an electrically programmable Ising network, which opens up new avenues to design nanomagnet-based logic devices and neuromorphic computers

cond-mat.mes-hall↗

Strong lateral exchange coupling and current-induced switching in single-layer ferrimagnetic films with patterned compensation temperature

Strong, adjustable magnetic couplings are of great importance to all devices based on magnetic materials. Controlling the coupling between adjacent regions of a single magnetic layer, however, is challenging. In this work, we demonstrate strong exchange-based coupling between arbitrarily shaped regions of a single ferrimagnetic layer. This is achieved by spatially patterning the compensation temperature of the ferrimagnet by either oxidation or He+ irradiation. The coupling originates at the lateral interface between regions with different compensation temperature and scales inversely with their width. We show that this coupling generates large lateral exchange coupling fields and we demonstrate its application to control the switching of magnetically compensated dots with an electric current.

cond-mat.mes-hall↗

Wide-spectrum optical synthetic aperture imaging via spatial intensity interferometry

High resolution imaging is achieved using increasingly larger apertures and successively shorter wavelengths. Optical aperture synthesis is an important high-resolution imaging technology used in astronomy. Conventional long baseline amplitude interferometry is susceptible to uncontrollable phase fluctuations, and the technical difficulty increases rapidly as the wavelength decreases. The intensity interferometry inspired by HBT experiment is essentially insensitive to phase fluctuations, but suffers from a narrow spectral bandwidth which results in a lack of detection sensitivity. In this study, we propose optical synthetic aperture imaging based on spatial intensity interferometry. This not only realizes diffraction-limited optical aperture synthesis in a single shot, but also enables imaging with a wide spectral bandwidth. And this method is insensitive to the optical path difference between the sub-apertures. Simulations and experiments present optical aperture synthesis diffraction-limited imaging through spatial intensity interferometry in a 100 $nm$ spectral width of visible light, whose maximum optical path difference between the sub-apertures reach $69.36λ$. This technique is expected to provide a solution for optical aperture synthesis over kilometer-long baselines at optical wavelengths.

physics.optics↗

Hyperspectral image reconstruction for spectral camera based on ghost imaging via sparsity constraints using V-DUnet

Spectral camera based on ghost imaging via sparsity constraints (GISC spectral camera) obtains three-dimensional (3D) hyperspectral information with two-dimensional (2D) compressive measurements in a single shot, which has attracted much attention in recent years. However, its imaging quality and real-time performance of reconstruction still need to be further improved. Recently, deep learning has shown great potential in improving the reconstruction quality and reconstruction speed for computational imaging. When applying deep learning into GISC spectral camera, there are several challenges need to be solved: 1) how to deal with the large amount of 3D hyperspectral data, 2) how to reduce the influence caused by the uncertainty of the random reference measurements, 3) how to improve the reconstructed image quality as far as possible. In this paper, we present an end-to-end V-DUnet for the reconstruction of 3D hyperspectral data in GISC spectral camera. To reduce the influence caused by the uncertainty of the measurement matrix and enhance the reconstructed image quality, both differential ghost imaging results and the detected measurements are sent into the network's inputs. Compared with compressive sensing algorithm, such as PICHCS and TwIST, it not only significantly improves the imaging quality with high noise immunity, but also speeds up the reconstruction time by more than two orders of magnitude.

eess.IV↗

Towards Top-Down Just Noticeable Difference Estimation of Natural Images

Just noticeable difference (JND) of natural images refers to the maximum pixel intensity change magnitude that typical human visual system (HVS) cannot perceive. Existing efforts on JND estimation mainly dedicate to modeling the diverse masking effects in either/both spatial or/and frequency domains, and then fusing them into an overall JND estimate. In this work, we turn to a dramatically different way to address this problem with a top-down design philosophy. Instead of explicitly formulating and fusing different masking effects in a bottom-up way, the proposed JND estimation model dedicates to first predicting a critical perceptual lossless (CPL) counterpart of the original image and then calculating the difference map between the original image and the predicted CPL image as the JND map. We conduct subjective experiments to determine the critical points of 500 images and find that the distribution of cumulative normalized KLT coefficient energy values over all 500 images at these critical points can be well characterized by a Weibull distribution. Given a testing image, its corresponding critical point is determined by a simple weighted average scheme where the weights are determined by a fitted Weibull distribution function. The performance of the proposed JND model is evaluated explicitly with direct JND prediction and implicitly with two applications including JND-guided noise injection and JND-guided image compression. Experimental results have demonstrated that our proposed JND model can achieve better performance than several latest JND models. In addition, we also compare the proposed JND model with existing visual difference predicator (VDP) metrics in terms of the capability in distortion detection and discrimination. The results indicate that our JND model also has a good performance in this task.

eess.IV↗

Engineering of Intrinsic Chiral Torques in Magnetic Thin Films Based on the Dzyaloshinskii-Moriya Interaction

The establishment of chiral coupling in thin magnetic films with inhomogeneous anisotropy has led to the development of artificial systems of fundamental and technological interest. The chiral coupling itself is enabled by the Dzyaloshinskii-Moriya interaction (DMI) enforced by the patterned noncollinear magnetization. Here, we create a domain wall track with out-of-plane magnetization coupled on each side to a narrow parallel strip with in-plane magnetization. With this we show that the chiral torques emerging from the DMI at the boundary between the regions of noncollinear magnetization in a single magnetic layer can be used to bias the domain wall velocity. To tune the chiral torques, the design of the magnetic racetracks can be modified by varying the width of the tracks or the width of the transition region between noncollinear magnetizations, reaching effective chiral magnetic fields of up to 7.8 mT. Furthermore, we show how the magnitude of the chiral torques can be estimated by measuring asymmetric domain wall velocities, and demonstrate spontaneous domain wall motion propelled by intrinsic torques even in the absence of any external driving force.

cond-mat.mes-hall↗

Deep learning tackles single-cell analysis A survey of deep learning for scRNA-seq analysis

Since its selection as the method of the year in 2013, single-cell technologies have become mature enough to provide answers to complex research questions. With the growth of single-cell profiling technologies, there has also been a significant increase in data collected from single-cell profilings, resulting in computational challenges to process these massive and complicated datasets. To address these challenges, deep learning (DL) is positioning as a competitive alternative for single-cell analyses besides the traditional machine learning approaches. Here we present a processing pipeline of single-cell RNA-seq data, survey a total of 25 DL algorithms and their applicability for a specific step in the processing pipeline. Specifically, we establish a unified mathematical representation of all variational autoencoder, autoencoder, and generative adversarial network models, compare the training strategies and loss functions for these models, and relate the loss functions of these models to specific objectives of the data processing step. Such presentation will allow readers to choose suitable algorithms for their particular objective at each step in the pipeline. We envision that this survey will serve as an important information portal for learning the application of DL for scRNA-seq analysis and inspire innovative use of DL to address a broader range of new challenges in emerging multi-omics and spatial single-cell sequencing.

q-bio.GN↗