arXiv · 2609.34550
Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning
Abstract
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide frame-level visual guidance for VLA fine-tuning. During training, recorded gaze locations are rendered as crosshairs on the robot's head-camera images. At deployment, a lightweight predictor estimates gaze locations from recent images and the instruction, supplying the same type of visual prompt without an eye tracker or changes to the policy architecture. Instantiated with $π_0$, gaze prompting increases mean success from $26.3\%$ to $56.0\%$ across six real-world bimanual manipulation tasks, with gains also observed when a single policy is trained on all six tasks. We release \textsc{GazeMani}, a dataset of $1{,}200$ teleoperated trajectories with synchronized gaze.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yihan Zhou, Rui Yan, Mingcong Li, Zheyuan Huang, Xu Yang, Xueyang Guo, Yilin Mo. 2026-09-28. Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning. https://arxiv.org/abs/2609.34550
Cite the original work for its findings. Save a collection to share your selection of sources.