arXiv · 2610.02943
Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy
Abstract
Multimodal 3D Human Pose Estimation (3D HPE) combines complementary information from RGB, LiDAR, and mmWave radar, but models trained on correlated observations from the same individuals, raise privacy risks overlooked by record level analysis. We present a unified framework for multimodal 3D HPE that couples kinematics-induced sensor fusion with subject level privacy auditing and private training. First, our multimodal model aligns modality specific joint representation, injects skeletal structure and adaptively aggregates complementary sensor evidence for accurate pose prediction. Second, we formulate a black-box subject membership inference attack for 3D HPE, complemented by an empirical pointwise maximal leakage analysis, which characterizes how individual attack score outcomes change inference about the membership outcome. Third, we instantiate user-level differential privacy via Action Temporal Stratification, a population weighted within-subject sampling strategy that enforces action and temporal coverage. We evaluate our framework on the MM-Fi dataset across three diverse experimental protocols. Source-code will be released upon acceptance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kaushik Bhargav Sivangi, Fani Deligianni. 2026-10-02. Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy. https://arxiv.org/abs/2610.02943
Cite the original work for its findings. Save a collection to share your selection of sources.