SearcharxivSearch

arXiv subjects

Jianhui Zhong

Publications and source records attributed to Jianhui Zhong.

10 recordsLinked to original sources

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought

Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Traditional methods typically encode input images into static feature representations and perform reasoning primarily within the linguistic domain, often overlooking the rich perceptual cues and explicit spatial geometry inherent in images. In this study, we aim to mitigate the cognitive vulnerability of models in interpreting gestural spatial relations by proposing PointVG-R, a reasoning-guided Multi-modal Large Language Model (MLLM). PointVG-R introduces geometric-aware reasoning for pointing-based grounding, enabling the model to think with images through the strategic integration of Reinforcement Learning (RL) and cold-start data. Specifically, we design a novel geometric reasoning pipeline that simulates the iterative cognitive process humans employ when interpreting pointing gestures. Furthermore, we construct EgoPoint-CoT, a high-quality visual Chain-of-Thought (CoT) dataset featuring detailed reasoning trajectories to guide the model via Supervised Fine-Tuning (SFT) and RL. To address the varying quality of learning signals encountered during training, we further propose an Adaptive Importance Weighting strategy based on Group Variance, which dynamically adjusts reward signals to optimize the learning process. Experimental results demonstrate that PointVG-R achieves SOTA performance, outperforming the baseline by $\textbf{15.86}$ points in mIoU. Extensive ablation studies further validate the efficacy of our proposed modules. Code: https://github.com/lingli1724/PointVG-R.

cs.CV

Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision

Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores non-verbal deictic cues prevalent in real-world interactions. In natural egocentric engagements, hand-pointing combined with speech forms the most intuitive referring mechanism. To bridge this gap, we introduce EgoPoint-Ground, the first large-scale multimodal dataset dedicated to egocentric deictic visual grounding. Comprising over \textbf{15k} interactive samples in complex scenes, the dataset provides rich, multi-grained annotations including hand-target bounding box pairs and dense semantic captions. We establish a comprehensive benchmark for hand-pointing referring expression resolution, evaluating a wide spectrum of mainstream Multimodal Large Language Models (MLLMs) and state-of-the-art VG architectures. Furthermore, we propose SV-CoT, a novel baseline framework that reformulates grounding as a structured inference process, synergizing gestural and linguistic cues through a Visual Chain-of-Thought paradigm. Extensive experiments demonstrate that SV-CoT achieves an $\textbf{11.7\%}$ absolute improvement over existing methods, effectively mitigating semantic ambiguity and advancing the capability of agents to comprehend multimodal physical intents. The dataset and code will be made publicly available.

cs.CV

High-resolution myelin-water fraction and quantitative relaxation mapping using 3D ViSTa-MR fingerprinting

Purpose: This study aims to develop a high-resolution whole-brain multi-parametric quantitative MRI approach for simultaneous mapping of myelin-water fraction (MWF), T1, T2, and proton-density (PD), all within a clinically feasible scan time. Methods: We developed 3D ViSTa-MRF, which combined Visualization of Short Transverse relaxation time component (ViSTa) technique with MR Fingerprinting (MRF), to achieve high-fidelity whole-brain MWF and T1/T2/PD mapping on a clinical 3T scanner. To achieve fast acquisition and memory-efficient reconstruction, the ViSTa-MRF sequence leverages an optimized 3D tiny-golden-angle-shuffling spiral-projection acquisition and joint spatial-temporal subspace reconstruction with optimized preconditioning algorithm. With the proposed ViSTa-MRF approach, high-fidelity direct MWF mapping was achieved without a need for multi-compartment fitting that could introduce bias and/or noise from additional assumptions or priors. Results: The in-vivo results demonstrate the effectiveness of the proposed acquisition and reconstruction framework to provide fast multi-parametric mapping with high SNR and good quality. The in-vivo results of 1mm- and 0.66mm-iso datasets indicate that the MWF values measured by the proposed method are consistent with standard ViSTa results that are 30x slower with lower SNR. Furthermore, we applied the proposed method to enable 5-minute whole-brain 1mm-iso assessment of MWF and T1/T2/PD mappings for infant brain development and for post-mortem brain samples. Conclusions: In this work, we have developed a 3D ViSTa-MRF technique that enables the acquisition of whole-brain MWF, quantitative T1, T2, and PD maps at 1mm and 0.66mm isotropic resolution in 5 and 15 minutes, respectively. This advancement allows for quantitative investigations of myelination changes in the brain.

physics.med-ph

Model-based Synthetic Data-driven Learning (MOST-DL): Application in Single-shot T2 Mapping with Severe Head Motion Using Overlapping-echo Acquisition

Use of synthetic data has provided a potential solution for addressing unavailable or insufficient training samples in deep learning-based magnetic resonance imaging (MRI). However, the challenge brought by domain gap between synthetic and real data is usually encountered, especially under complex experimental conditions. In this study, by combining Bloch simulation and general MRI models, we propose a framework for addressing the lack of training data in supervised learning scenarios, termed MOST-DL. A challenging application is demonstrated to verify the proposed framework and achieve motion-robust T2 mapping using single-shot overlapping-echo acquisition. We decompose the process into two main steps: (1) calibrationless parallel reconstruction for ultra-fast pulse sequence and (2) intra-shot motion correction for T2 mapping. To bridge the domain gap, realistic textures from a public database and various imperfection simulations were explored. The neural network was first trained with pure synthetic data and then evaluated with in vivo human brain. Both simulation and in vivo experiments show that the MOST-DL method significantly reduces ghosting and motion artifacts in T2 maps in the presence of unpredictable subject movement and has the potential to be applied to motion-prone patients in the clinic.

eess.IV

Optimized multi-axis spiral projection MR fingerprinting with subspace reconstruction for rapid whole-brain high-isotropic-resolution quantitative imaging

Purpose: To improve image quality and accelerate the acquisition of 3D MRF. Methods: Building on the multi-axis spiral-projection MRF technique, a subspace reconstruction with locally low rank (LLR) constraint and a modified spiral-projection spatiotemporal encoding scheme termed tiny-golden-angle-shuffling (TGAS) were implemented for rapid whole-brain high-resolution quantitative mapping. The LLR regularization parameter and the number of subspace bases were tuned using retrospective in-vivo data and simulated examinations, respectively. B0 inhomogeneity correction using multi-frequency interpolation was incorporated into the subspace reconstruction to further improve the image quality by mitigating blurring caused by off-resonance effect. Results: The proposed MRF acquisition and reconstruction framework can produce provide high quality 1-mm isotropic whole-brain quantitative maps in a total acquisition time of 1 minute 55 seconds, with higher-quality results than ones obtained from the previous approach in 6 minutes. The comparison of quantitative results indicates that neither the subspace reconstruction nor the TGAS trajectory induce bias for T1 and T2 mapping. High quality whole-brain MRF data were also obtained at 0.66-mm isotropic resolution in 4 minutes using the proposed technique, where the increased resolution was shown to improve visualization of subtle brain structures. Conclusion: The proposed TGAS-SPI-MRF with optimized spiral-projection trajectory and subspace reconstruction can enable high-resolution quantitative mapping with faster acquisition speed.

physics.med-ph

Single-Shell NODDI Using Dictionary Learner Estimated Isotropic Volume Fraction

Neurite orientation dispersion and density imaging (NODDI) enables the assessment of intracellular, extracellular and free water signals from multi-shell diffusion MRI data. It is an insightful approach to characterize brain tissue microstructure. Single-shell reconstruction for NODDI parameters has been discouraged in previous studies caused by failure when fitting, especially for the neurite density index (NDI). Here, we investigated the possibility of creating robust NODDI parameter maps with single-shell data, using the isotropic volume fraction (fISO) as prior. Prior estimation was made independent of the NODDI model constraint using a dictionary learning approach. First, we used a stochastic sparse dictionary-based network (DictNet) in predicting fISO which is trained with data obtained from in vivo and simulated diffusion MRI data. In single-shell cases, the mean diffusivity (MD) and raw T2 signal with no diffusion weighting (S0) was incorporated in the dictionary for the fISO estimation. Then, the NODDI framework was used with the known fISO to estimate the NDI and orientation dispersion index (ODI). The fISO estimated by our model was compared with other fISO estimators in the simulation. Further, using both synthetic data simulation and human data collected on a 3T scanner, we compared the performance of our dictionary-based learning prior NODDI (DLpN) with the original NODDI for both single-shell and multi-shell data. Our results suggest that DLpN derived NDI and ODI parameters for single-shell protocols are comparable with original multi-shell NODDI, and protocol with b=2000 s/mm2 performs the best (error ~5% in white and grey matter). This may allow NODDI evaluation of studies on single-shell data by multi-shell scanning of two subjects for DictNet fISO training.

physics.med-ph

Efficient T2 mapping with Blip-up/down EPI and gSlider-SMS (T2-BUDA-gSlider)

Purpose: To rapidly obtain high isotropic-resolution T2 maps with whole-brain coverage and high geometric fidelity. Methods: A T2 blip-up/down echo planar imaging (EPI) acquisition with generalized Slice-dithered enhanced resolution (T2-BUDA-gSlider) is proposed. A radiofrequency (RF)-encoded multi-slab spin-echo EPI acquisition with multiple echo times (TEs) was developed to obtain high SNR efficiency with reduced repetition time (TR). This was combined with an interleaved 2-shot EPI acquisition using blip-up/down phase encoding. An estimated field map was incorporated into the joint multi-shot EPI reconstruction with a structured low rank constraint to achieve distortion-free and robust reconstruction for each slab without navigation. A Bloch simulated subspace model was integrated into gSlider reconstruction and utilized for T2 quantification. Results: In vivo results demonstrated that the T2 values estimated by the proposed method were consistent with gold standard spin-echo acquisition. Compared to the reference 3D fast spin echo (FSE) images, distortion caused by off-resonance and eddy current effects were effectively mitigated. Conclusion: BUDA-gSlider SE-EPI acquisition and gSlider-subspace joint reconstruction enabled distortion-free whole-brain T2 mapping in 2 min at ~1 mm3 isotropic resolution, which could bring significant benefits to related clinical and neuroscience applications.

physics.med-ph

Robust diffusion parametric mapping of motion-corrupted data with a three-dimensional convolutional neural network

Head motion is inevitable in the acquisition of diffusion-weighted images, especially for certain motion-prone subjects and for data gathering of advanced diffusion models with prolonged scan times. Deficient accuracy of motion correction cause deterioration in the quality of diffusion model reconstruction, thus affecting the derived measures. This results in either loss of data, or introducing bias in outcomes from data of different motion levels, or both. Hence minimizing motion effects and reutilizing motion-contaminated data becomes vital to quantitative studies. We have previously developed a 3-dimensional hierarchical convolution neural network (3D H-CNN) for robust diffusion kurtosis mapping from under-sampled data. In this study, we propose to extend this method to motion-contaminated data for robust recovery of diffusion model-derived measures with a process of motion assessment and corrupted volume rejection. We validate the proposed pipeline in two in-vivo datasets. Results from the first dataset of individual subjects show that all the diffusion tensor and kurtosis tensor-derived measures from the new pipeline are minimally sensitive to motion effects, and are comparable to the motion-free reference with as few as eight volumes retained from the motion-contaminated data. Results from the second dataset of a group of children with attention deficit hyperactivity disorder demonstrate the ability of our approach in ameliorating spurious group differences due to head motion. This method shows great potential for exploiting some valuable but motion-corrupted DWI data which are likely to be discarded otherwise, and applying to data with different motion level thus improving their utilization and statistic power.

physics.med-ph

Ultrashort Echo Time Magnetic Resonance Fingerprinting (UTE-MRF) for Simultaneous Quantification of Long and Ultrashort T2 Tissues

Purpose: To demonstrate an ultrashort echo time magnetic resonance fingerprinting (UTE-MRF) method that can simultaneously quantify tissue relaxometries for muscle and bone in musculoskeletal systems and tissue components in brain and therefore can synthesize pseudo-CT images. Methods: A FISP-MRF sequence with half pulse excitation and half spoke radial acquisition was designed to sample fast T2 decay signals. Sinusoidal echo time (TE) pattern was applied to enhance MRF sensitivity for tissues with short and ultrashort T2 values. The performance of UTE-MRF was evaluated via simulations, phantoms, and in vivo experiments. Results: A minimal TE of 0.05 ms was achieved in UTE-MRF. Simulations indicated that extension of TE sampling increased T2 quantification accuracy in cortical bone and tendon, and had little impact on long T2 muscle quantifications. For a rubber phantom, an average T1/T2 of 162/1.07 ms from UTE-MRF were compared well with gold standard T2 of 190 ms from IR-UTE and T2* of 1.03 ms from UTE sequence. For a long T2 agarose phantom, the linear regression slope between UTE-MRF and gold standard was 1.07 (R2=0.991) for T1 and 1.04 (R2=0.994) for T2. In vivo experiments showed the detection of cortical bone and Achilles tendon, where the averaged T2 was respectively 1.0 ms and 15 ms. Scalp images were in good agreement with CT. Conclusion: UTE-MRF with sinusoidal TE variations shows its capability to produce pseudo-CT images and simultaneously output T1, T2, proton density, and B0 maps for tissues with long T2 and short/ultrashort T2 in the brain and musculoskeletal system.

physics.med-ph

High Efficient Reconstruction of Single-shot T2 Mapping from OverLapping-Echo Detachment Planar Imaging Based on Deep Residual Network

Purpose: An end-to-end deep convolutional neural network (CNN) based on deep residual network (ResNet) was proposed to efficiently reconstruct reliable T2 mapping from single-shot OverLapping-Echo Detachment (OLED) planar imaging. Methods: The training dataset was obtained from simulations carried out on SPROM software developed by our group. The relationship between the original OLED image containing two echo signals and the corresponded T2 mapping was learned by ResNet training. After the ResNet was trained, it was applied to reconstruct the T2 mapping from simulation and in vivo human brain data. Results: Though the ResNet was trained entirely on simulated data, the trained network was generalized well to real human brain data. The results from simulation and in vivo human brain experiments show that the proposed method significantly outperformed the echo-detachment-based method. Reliable T2 mapping was achieved within tens of milliseconds after the network had been trained while the echo-detachment-based OLED reconstruction method took minutes. Conclusion: The proposed method will greatly facilitate real-time dynamic and quantitative MR imaging via OLED sequence, and ResNet has the potential to reconstruct images from complex MRI sequence efficiently.

cs.CV