arXiv · 2608.04798
Reference-Based Manipulation: A Framework and Pipeline for Multimodal Spatial Reasoning
Abstract
When manipulating objects in immersive platforms through speech and gesture, users naturally construct spatial references, referring to scene entities, their bodies, or the environment. Leveraging spatial cognition theories, this work systematically examines how users construct and communicate spatial intent. Using a custom toolkit, we conducted a Wizard-of-Oz study to observe unconstrained multimodal (speech + gesture) input patterns in Virtual Reality for scene construction. Based on these findings, we formalize a framework that decomposes spatial references into three core components: Source, Anchor, and Frame, while characterizing their compositional strategies and explicitness. We demonstrate the utility of this Reference-based Manipulation framework by implementing an LLM-based pipeline featuring a set of example interaction techniques with a preliminary technical evaluation. Finally, we discuss key lessons learned for supporting reference-based spatial interaction.
Explore related subjects
Keep this discovery
Yangyang He, Zhuangze Hou, Yonglin Chen, Can Liu. 2026-08-05. Reference-Based Manipulation: A Framework and Pipeline for Multimodal Spatial Reasoning. https://doi.org/10.1145/3830398.3830657
Cite the original work for its findings. Save a collection to share your selection of sources.