SearcharxivSearch

arXiv subjects

Haodong Xu

Publications and source records attributed to Haodong Xu.

7 recordsLinked to original sources

PAS-Mamba: Phase-Amplitude-Spatial State Space Model for MRI Reconstruction

Joint feature modeling in both the spatial and frequency domains has become a mainstream approach in MRI reconstruction. However, existing methods generally treat the frequency domain as a whole, neglecting the differences in the information carried by its internal components. According to Fourier transform theory, phase and amplitude represent different types of information in the image. Our spectrum swapping experiments show that magnitude mainly reflects pixel-level intensity, while phase predominantly governs image structure. To prevent interference between phase and magnitude feature learning caused by unified frequency-domain modeling, we propose the Phase-Amplitude-Spatial State Space Model (PAS-Mamba) for MRI Reconstruction, a framework that decouples phase and magnitude modeling in the frequency domain and combines it with image-domain features for better reconstruction. In the image domain, LocalMamba preserves spatial locality to sharpen fine anatomical details. In frequency domain, we disentangle amplitude and phase into two specialized branches to avoid representational coupling. To respect the concentric geometry of frequency information, we propose Circular Frequency Domain Scanning (CFDS) to serialize features from low to high frequencies. Finally, a Dual-Domain Complementary Fusion Module (DDCFM) adaptively fuses amplitude phase representations and enables bidirectional exchange between frequency and image domains, delivering superior reconstruction. Extensive experiments on the IXI and fastMRI knee datasets show that PAS-Mamba consistently outperforms state of the art reconstruction methods.

cs.CV

Modulation Instability-Induced Multimode Squeezing in Quadratic Frequency Combs

Lithium niobate (LN) microring resonators, characterized by an exceptionally high second-order nonlinear coefficient and superior electro-optic tunability, serve as an outstanding platform for the precise control of integrated quantum frequency combs (QFCs). In this study, we introduce a bipartite entanglement criterion to investigate the pairwise entanglement characteristics of QFCs generated via the spontaneous parametric down-conversion (SPDC) process in lithium niobate microring resonators operating below threshold. Furthermore, we propose a universal framework for analyzing multimode squeezing in quadratic frequency combs, enabling the realization of ultrabroadband and high-degree multimode squeezing. We further reveal the underlying physical mechanism: modulation instability (MI), regulated by temporal walk-off control, not only enables the formation of frequency combs but also induces multimode squeezing in the corresponding resonant modes. This study uncovers the previously unexplored role of on-chip multimode squeezing in quadratic frequency combs while facilitating collective noise suppression across multiple modes, thus holding substantial potential for advancing quantum precision measurement and quantum information processing.

quant-ph

A Review of Behavioral Closed-Loop Paradigm from Sensing to Intervention for Ingestion Health

Ingestive behavior plays a critical role in health, yet many existing interventions remain limited to static guidance or manual self-tracking. With the increasing integration of sensors, context-aware computing, and perceptual computing, recent systems have begun to support closed-loop interventions that dynamically sense user behavior and provide feedback during or around ingestion episodes. In this survey, we review 136 studies that leverage sensor-enabled or interaction-mediated approaches to influence ingestive behavior. We propose a behavioral closed-loop paradigm rooted in context-aware computing and inspired by HCI behavior change frameworks, comprising four components: target behaviors, sensing modalities, reasoning and intervention strategies. A taxonomy of sensing and intervention modalities is presented, organized along human- and environment-based dimensions. Our analysis also examines evaluation methods and design trends across different modality-behavior pairings. This review reveals prevailing patterns and critical gaps, offering design insights for future adaptive and context-aware ingestion health interventions.

cs.HC

SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context

Endeavors have been made to explore Large Language Models for video analysis (Video-LLMs), particularly in understanding and interpreting long videos. However, existing Video-LLMs still face challenges in effectively integrating the rich and diverse audio-visual information inherent in long videos, which is crucial for comprehensive understanding. This raises the question: how can we leverage embedded audio-visual information to enhance long video understanding? Therefore, (i) we introduce SAVEn-Vid, the first-ever long audio-visual video dataset comprising over 58k audio-visual instructions. (ii) From the model perspective, we propose a time-aware Audio-Visual Large Language Model (AV-LLM), SAVEnVideo, fine-tuned on SAVEn-Vid. (iii) Besides, we present AVBench, a benchmark containing 2,500 QAs designed to evaluate models on enhanced audio-visual comprehension tasks within long video, challenging their ability to handle intricate audio-visual interactions. Experiments on AVBench reveal the limitations of current AV-LLMs. Experiments also demonstrate that SAVEnVideo outperforms the best Video-LLM by 3.61% on the zero-shot long video task (Video-MME) and surpasses the leading audio-visual LLM by 1.29% on the zero-shot audio-visual task (Music-AVQA). Consequently, at the 7B parameter scale, SAVEnVideo can achieve state-of-the-art performance. Our dataset and code will be released at https://ljungang.github.io/SAVEn-Vid/ upon acceptance.

cs.CV

Frequency-dependent squeezing via Einstein-Podolsky-Rosen entanglement based on silicon nitride microring resonators

Significant efforts have been made to enhance the performance of displacement sensors limited by quantum noise, such as gravitational wave detectors. Techniques like frequency-dependent squeezing have overcome the standard quantum limit in optomechanical force measurements, leading to substantial overall progress. These advancements, coupled with major developments in integrated photonics, have paved the way for the emergence of integrated Kerr quantum frequency combs (QFCs). A platform has been established for designing EPR entangled quantum frequency combs using on-chip silicon nitride microring resonators, enabling thorough analysis and optimization of entanglement performance, as well as effective noise reduction adjustments. This platform, incorporating the quantum dynamics of Kerr nonlinear microresonators, supports at least 12 continuous-variable quantum modes in the form of 6 simultaneous two-mode squeezed pairs (EPR entangled pairs). Additionally, by selecting the detection angle of the idler mode, a single-mode squeezed state is generated in the signal mode. Given the frequency-dependent nature of the detection angle, frequency-dependent squeezing is achieved. A comparative analysis of the results under different dispersion conditions is also conducted.

quant-ph

Pedestrian Recognition with Radar Data-Enhanced Deep Learning Approach Based on Micro-Doppler Signatures

As a hot topic in recent years, the ability of pedestrians identification based on radar micro-Doppler signatures is limited by the lack of adequate training data. In this paper, we propose a data-enhanced multi-characteristic learning (DEMCL) model with data enhancement (DE) module and multi-characteristic learning (MCL) module to learn more complementary pedestrian micro-Doppler (m-D) signatures. In DE module, a range-Doppler generative adversarial network (RDGAN) is proposed to enhance free walking datasets, and MCL module with multi-scale convolution neural network (MCNN) and radial basis function neural network (RBFNN) is trained to learn m-D signatures extracted from enhanced datasets. Experimental results show that our model is 3.33% to 10.24% more accurate than other studies and has a short run time of 0.9324 seconds on a 25-minute walking dataset.

eess.SP

A Multi-Characteristic Learning Method with Micro-Doppler Signatures for Pedestrian Identification

The identification of pedestrians using radar micro-Doppler signatures has become a hot topic in recent years. In this paper, we propose a multi-characteristic learning (MCL) model with clusters to jointly learn discrepant pedestrian micro-Doppler signatures and fuse the knowledge learned from each cluster into final decisions. Time-Doppler spectrogram (TDS) and signal statistical features extracted from FMCW radar, as two categories of micro-Doppler signatures, are used in MCL to learn the micro-motion information inside pedestrians' free walking patterns. The experimental results show that our model achieves a higher accuracy rate and is more stable for pedestrian identification than other studies, which make our model more practical.

eess.SP