arXiv · 2609.25750
Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera
Abstract
Manipulation requires both broad scene awareness and detailed local feedback, yet conventional camera rigs provide them through separate front and wrist cameras. We present Fisheye-VLA, a visual interface that brings these capabilities together using a single passive fisheye. A global view preserves the workspace, while local perspective crops direct detail toward the interaction. The key design question is where this local visual budget should go. We answer it through a controlled re-rendering study, comparing alternative crop directions on the same recorded observations. The study finds that end-effector-centered views capture most of the estimated benefit of a much larger candidate pool, motivating a compact allocation around both hands. Our interface uses calibrated end-effector projection and motion lead to track the crops, while a shared ray encoding preserves their spatial meaning as they move. Integrated with a pretrained VLA, it achieves 84% and 82% success in the two expanded tabletop regions, where some target placements extend beyond the front-camera coverage, and supports shelf and conveyor manipulation. Ablations show that local crops and their viewing directions become more important in the larger workspace regions. The results demonstrate that a single fisheye can support these manipulation tasks without physical wrist cameras.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ziang Ren, Zike Yan, Raymond Zhang, Xuguo He, Zhongyu Li. 2026-09-22. Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera. https://arxiv.org/abs/2609.25750
Cite the original work for its findings. Save a collection to share your selection of sources.