SearcharxivSearch

arXiv subjects

Yifan Zheng

Publications and source records attributed to Yifan Zheng.

6 recordsLinked to original sources

A Positron Range Correction with Texture Preservation Framework in PET Imaging

Positron range (PR) blurring is a fundamental resolution limitation in PET imaging with high-energy positron emitters such as 82Rb, causing contrast loss and spill-out effects across heterogeneous tissue interfaces. We propose PRC-TP, a positron range correction (PRC) framework with explicit texture preservation that decouples deterministic resolution recovery from stochastic texture restoration. A nnFormer-based neural network (NN) was trained on patient-derived Monte Carlo simulations to map PR-degraded 82Rb reconstructions to PR-free references using attenuation maps as anatomical context. However, this NN also significantly removed the noise in the images, which could impact some texture analysis methods or make the images look unrealistic. An auxiliary Noise2Noise model estimates that smoothing effect, enabling texture extraction and transfer to the PR-corrected prediction through Model-consistent Texture Re-Injection (MTRI). In simulated patients, PRC-TP preserved contrast recovery close to ground truth (GT) (98.96-99.04%) while restoring noise and CNR closer to the reference. The function-based MTRI formulation achieved near unity global texture amplitude agreement with GT (0.997 +/- 0.011), reducing the input texture amplitude bias (0.951 +/- 0.011). Radiomics analysis showed improved agreement with GT across texture-sensitive feature families. A clinical 82Rb evaluation showed trends consistent with simulations, including comparable contrast-ratio increase (10.18% vs. 10.99%) and restoration of texture suppressed by PRC. These results support PRC-TP as a practical framework for resolution recovery with acquisition-consistent texture preservation in PET imaging. Submitted to IEEE TRPMS.

physics.med-ph

Tantalizing Evidence of Reionization Relics in the eBOSS DR16 Ly$\boldsymbolα$ Forest Correlations: a Preference for Early Reionization

Cosmic reionization of HI leaves enduring relics in the post-reionization intergalactic medium, potentially influencing the Lyman-$α$ (Ly$α$) forest down to redshifts as low as $z \approx 2$, which is the so-called ''memory of reionization'' effect. Here, we re-analyze the baryonic acoustic oscillation (BAO) measurements from Ly$α$ absorption and quasar correlations using data from the extended Baryonic Oscillation Spectroscopic Survey (eBOSS) Data Release 16 (DR16), incorporating for the first time the memory of reionization in the Ly$α$ forest. Three distinct scenarios of reionization timeline are considered in our analyses. We find that the recovered BAO parameters ($α_\parallel$, $α_\perp$) remain consistent with the original eBOSS DR16 analysis. However, models incorporating reionization relics provide a better fit to the data, with a tantalizing preference for early reionization, consistent with recent findings from the James Webb Space Telescope. Furthermore, the inclusion of reionization relics significantly impacts the non-BAO parameters. For instance, we report deviations of up to $3σ$ in the Ly$α$ redshift-space distortion parameter and $\sim7σ$ in the linear Ly$α$ bias for the late reionization scenario. Our findings suggest that the eBOSS Ly$α$ data is more accurately described by models that incorporate a broadband enhancement to the Ly$α$ forest power spectrum, highlighting the importance of accounting for reionization relics in cosmological analyses.

astro-ph.CO

PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation

The fragmentation between high-level task semantics and low-level geometric features remains a persistent challenge in robotic manipulation. While vision-language models (VLMs) have shown promise in generating affordance-aware visual representations, the lack of semantic grounding in canonical spaces and reliance on manual annotations severely limit their ability to capture dynamic semantic-affordance relationships. To address these, we propose Primitive-Aware Semantic Grounding (PASG), a closed-loop framework that introduces: (1) Automatic primitive extraction through geometric feature aggregation, enabling cross-category detection of keypoints and axes; (2) VLM-driven semantic anchoring that dynamically couples geometric primitives with functional affordances and task-relevant description; (3) A spatial-semantic reasoning benchmark and a fine-tuned VLM (Qwen2.5VL-PA). We demonstrate PASG's effectiveness in practical robotic manipulation tasks across diverse scenarios, achieving performance comparable to manual annotations. PASG achieves a finer-grained semantic-affordance understanding of objects, establishing a unified paradigm for bridging geometric primitives with task semantics in robotic manipulation.

cs.CV

SafeEar: Content Privacy-Preserving Audio Deepfake Detection

Text-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals. Existing countermeasures largely focus on determining the genuineness of speech based on complete original audio recordings, which however often contain private content. This oversight may refrain deepfake detection from many applications, particularly in scenarios involving sensitive information like business secrets. In this paper, we propose SafeEar, a novel framework that aims to detect deepfake audios without relying on accessing the speech content within. Our key idea is to devise a neural audio codec into a novel decoupling model that well separates the semantic and acoustic information from audio samples, and only use the acoustic information (e.g., prosody and timbre) for deepfake detection. In this way, no semantic content will be exposed to the detector. To overcome the challenge of identifying diverse deepfake audio without semantic clues, we enhance our deepfake detector with real-world codec augmentation. Extensive experiments conducted on four benchmark datasets demonstrate SafeEar's effectiveness in detecting various deepfake techniques with an equal error rate (EER) down to 2.02%. Simultaneously, it shields five-language speech content from being deciphered by both machine and human auditory analysis, demonstrated by word error rates (WERs) all above 93.93% and our user study. Furthermore, our benchmark constructed for anti-deepfake and anti-content recovery evaluation helps provide a basis for future research in the realms of audio privacy preservation and deepfake detection.

cs.CR

Deep Neural Networks in Video Human Action Recognition: A Review

Currently, video behavior recognition is one of the most foundational tasks of computer vision. The 2D neural networks of deep learning are built for recognizing pixel-level information such as images with RGB, RGB-D, or optical flow formats, with the current increasingly wide usage of surveillance video and more tasks related to human action recognition. There are increasing tasks requiring temporal information for frames dependency analysis. The researchers have widely studied video-based recognition rather than image-based(pixel-based) only to extract more informative elements from geometry tasks. Our current related research addresses multiple novel proposed research works and compares their advantages and disadvantages between the derived deep learning frameworks rather than machine learning frameworks. The comparison happened between existing frameworks and datasets, which are video format data only. Due to the specific properties of human actions and the increasingly wide usage of deep neural networks, we collected all research works within the last three years between 2020 to 2022. In our article, the performance of deep neural networks surpassed most of the techniques in the feature learning and extraction tasks, especially video action recognition.

cs.CV

Robustness of Optimal Energy Thresholds in Photon-counting Spectral CT

An important question when developing photon-counting detectors for computed tomography is how to select energy thresholds. In this work thresholds are optimized by maximizing signal-difference-to-noise ratio squared (SDNR2) in an optimally weighted image and signal-to-noise ratio squared (SNR2) in a gadolinium basis image in a silicon-strip detector and a cadmium zinc telluride (CZT) detector, factoring in pileup and imperfect energy response in both detectors. To investigate to what extent one single set of thresholds could be applied in various imaging tasks, the robustness of optimal thresholds with 2 to 8 bins is examined with the variation of phantom thicknesses and target materials. In contrast to previous studies, the optimal threshold locations don't always increase with increasing attenuation if pileup is included. Optimizing the thresholds for a 30 cm phantom yields near-optimal SDNR2 or SNR2 regardless of target tissue types and surrounding attenuation for both detectors. Having more than 3 bins reduces the need for changing the thresholds depending on anatomies and tissues. Using around 6 bins or 8 bins may give near-optimal SDNR2 or SNR2 without generating an unnecessarily large amount of data.

physics.med-ph