SearcharxivSearch

arXiv subjects

Xijia Wei

Publications and source records attributed to Xijia Wei.

3 recordsLinked to original sources

MAEPose: Self-Supervised Spatiotemporal Learning for Human Pose Estimation on mmWave Video

Millimetre-wave (mmWave) radar offers a more privacy-preserving alternative to RGB-based human pose estimation. However, existing methods typically rely on pre-extracted intermediate representations such as sparse point clouds or spectrogram images, where the rich spatiotemporal information naturally present in radar video streams is discarded for model learning, while such signal processing adds system complexity. In addition, existing solutions are mainly conducted in an end-to-end supervised manner without leveraging unlabelled raw video streams to learn generalized representations. In this study, we present MAEPose, a masked autoencoding-based human pose estimation approach that operates directly on mmWave spectrogram videos. MAEPose learns spatiotemporal motion-aware generalized representations from unlabelled radar video, and leverages its heatmap decoder for multi-frame pose estimation predictions. We evaluate it across three datasets based on leave-one-person-out cross-validation with rigorous statistical testing. MAEPose consistently outperforms state-of-the-art baselines by up to 22.1% in MPJPE p<0.05, and maintains robust accuracy under zero-shot bystander interference with only a 6.5% error increase. Ablation studies confirm that both the pre-training and the heatmap decoder contribute substantially, while modality analysis indicates that leveraging Range-Doppler video as input achieves better pose estimation performance than Range-Azimuth or their fusion, with lower computational cost.

cs.CV

Listening to the Mind: Earable Acoustic Sensing of Cognitive Load

Earable acoustic sensing offers a powerful and non-invasive modality for capturing fine-grained auditory and physiological signals directly from the ear canal, enabling continuous and context-aware monitoring of cognitive states. As earable devices become increasingly embedded in daily life, they provide a unique opportunity to sense mental effort and perceptual load in real time through auditory interactions. In this study, we present the first investigation of cognitive load inference through auditory perception using acoustic signals captured by off-the-shelf in-ear devices. We designed speech-based listening tasks to induce varying levels of cognitive load, while concurrently embedding acoustic stimuli to evoke Stimulus Frequency Otoacoustic Emission (SFOAEs) as a proxy for cochlear responsiveness. Statistical analysis revealed a significant association (p < 0.01) between increased cognitive load and changes in auditory sensitivity, with 63.2% of participants showing peak sensitivity at 3 kHz. Notably, sensitivity patterns also varied across demographic subgroups, suggesting opportunities for personalized sensing. Our findings demonstrate that earable acoustic sensing can support scalable, real-time cognitive load monitoring in natural settings, laying a foundation for future applications in augmented cognition, where everyday auditory technologies adapt to and support the users mental health.

cs.HC

CubeDN: Real-time Drone Detection in 3D Space from Dual mmWave Radar Cubes

As drone use has become more widespread, there is a critical need to ensure safety and security. A key element of this is robust and accurate drone detection and localization. While cameras and other optical sensors like LiDAR are commonly used for object detection, their performance degrades under adverse lighting and environmental conditions. Therefore, this has generated interest in finding more reliable alternatives, such as millimeter-wave (mmWave) radar. Recent research on mmWave radar object detection has predominantly focused on 2D detection of road users. Although these systems demonstrate excellent performance for 2D problems, they lack the sensing capability to measure elevation, which is essential for 3D drone detection. To address this gap, we propose CubeDN, a single-stage end-to-end radar object detection network specifically designed for flying drones. CubeDN overcomes challenges such as poor elevation resolution by utilizing a dual radar configuration and a novel deep learning pipeline. It simultaneously detects, localizes, and classifies drones of two sizes, achieving decimeter-level tracking accuracy at closer ranges with overall $95\%$ average precision (AP) and $85\%$ average recall (AR). Furthermore, CubeDN completes data processing and inference at 10Hz, making it highly suitable for practical applications.

cs.RO