arXiv · 2609.24424
Estimating Accurate Hand Pose in Camera Space with Vision Transformer
Abstract
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect of hand local poses and global wrist positions in the perspective projections. In particular, this coupling reflects that the projections are jointly determined by local hand poses, wrist positions, and camera intrinsics. To overcome these challenges, our framework proposes two key innovations: Transformation-Isomorphism Supervision for hand-depth information extraction and Perspective Information Embedding for resolving above coupling effect of local pose and wrist position, both integrated within the mainstream encoder-decoder architecture. Besides, we propose a novel framerate-aware multi-dataset training strategy for sequential pose refinement. Our fully integrated approach achieves at most 37.1\% superiority in CS-MJE over SOTA on HO3D. Project page: https://github.com/Mine268/CS-ViT.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kaiwen Ren, Yiran Jiang, Yongjing Ye, Shihong Xia. 2026-09-21. Estimating Accurate Hand Pose in Camera Space with Vision Transformer. https://doi.org/10.1145/3767308.3836423
Cite the original work for its findings. Save a collection to share your selection of sources.