arXiv · 2610.03105
A Benchmark for Spatially Grounded Gesture Generation
Abstract
Communication in shared space interweaves verbal and non-verbal signals, and pointing gestures anchor language to the environment: "put the cup on that one" is uninterpretable without the gesture that fixes the referent. Yet no common framework exists for evaluating whether generated gestures indicate their intended referent; distributional metrics reward a gesture aimed at the wrong object as long as it looks natural. We introduce a benchmark for spatially grounded gesture generation, comprising ~2K pointing-annotated clips from naturalistic VR dialogue with ground-truth 3D referents, a task in which systems must decide when, how and where to point within conversational speech, and a protocol that separates temporal alignment, spatial grounding and perceived naturalness. We also provide a flow-matching baseline, MM-Conv-Flow. Evaluating it alongside an independent retrieval-based system and captured human motion, we find that geometric grounding can exceed that of human pointing without any gain in perceived naturalness, showing that referential gesture quality must be measured along separate dimensions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Anna Deichler, Rishabh Dabral, Fethiye Irmak Dogan, Anindita Ghosh, Jonas Beskow. 2026-10-02. A Benchmark for Spatially Grounded Gesture Generation. https://arxiv.org/abs/2610.03105
Cite the original work for its findings. Save a collection to share your selection of sources.