SearcharxivSearch

arXiv subjects

Qidi Luo

Publications and source records attributed to Qidi Luo.

3 recordsLinked to original sources

Two-stage Respiratory Motion-resolved Radial MR Image Reconstruction Using an Interpretable Deep Unrolled Network

Due to the prolonged MRI encoding process, respiratory motion can cause undesired artifacts and image blurring, degrading image quality and limiting clinical applications in abdominal and pulmonary imaging. In this work, we develop a two-stage respiratory motion-resolved radial MR image reconstruction pipeline using an interpretable deep unrolled network (MoraNet), enabling high-quality imaging under free-breathing conditions. Firstly, low-resolution images are reconstructed from the central region of successive golden-angle radial k-space to extract respiratory motion signals. The binned k-space data based on the respiratory signal are then used to reconstruct the motion-resolved high-resolution image for each motion state. The MoraNet applies nonuniform fast Fourier transform (NUFFT) to operate radial encoding and convolutional neural network (CNN) modules to conduct image regularizations. The MoraNet was trained on retrospectively acquired lung MRI images for both fully sampled and undersampled acquisitions. The performance of the proposed method was evaluated on digital CT/MRI breathing XCAT (CoMBAT) phantom data, QUASAR motion phantom data acquired from a 1.0T MRI scanner and volunteer chest data acquired from a 1.5T MRI scanner. The MoraNet pipeline was compared with motion-averaged reconstruction and a conventional compressed sensing (CS)-based method in terms of SSIM, RMSE and computation time. Simulation and experimental results demonstrated that the proposed network could provide accurate respiratory signal estimation and enable effective motion correction. Compared with the CS method, the MoraNet preserved better structural details with lower RMSE and higher SSIM values at acceleration factor of 4, and meanwhile took ten-fold faster inference time.

physics.med-ph

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks (4x more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios (31 diverse scenarios), and thorough evaluation metrics, with 10,000 human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with 1,500 manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below 50 (100 in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The project website is at: https://99franklin.github.io/ocrbench_v2/

cs.CV

Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering

The development of Urdu scene text detection, recognition, and Visual Question Answering (VQA) technologies is crucial for advancing accessibility, information retrieval, and linguistic diversity in digital content, facilitating better understanding and interaction with Urdu-language visual data. This initiative seeks to bridge the gap between textual and visual comprehension. We propose a new multi-task Urdu scene text dataset comprising over 1000 natural scene images, which can be used for text detection, recognition, and VQA tasks. We provide fine-grained annotations for text instances, addressing the limitations of previous datasets for facing arbitrary-shaped texts. By incorporating additional annotation points, this dataset facilitates the development and assessment of methods that can handle diverse text layouts, intricate shapes, and non-standard orientations commonly encountered in real-world scenarios. Besides, the VQA annotations make it the first benchmark for the Urdu Text VQA method, which can prompt the development of Urdu scene text understanding. The proposed dataset is available at: https://github.com/Hiba-MeiRuan/Urdu-VQA-Dataset-/tree/main

cs.CV