arXiv · 2506.02556
Sign Language: Towards Sign Understanding for Robot Autonomy
Abstract
Navigational signs are common aids for human wayfinding and scene understanding, but are underutilized by robots. We argue that they benefit robot navigation and scene understanding, by directly encoding privileged information on actions, spatial regions, and relations. Interpreting signs in open-world settings remains a challenge owing to the complexity of scenes and signs, but recent advances in vision-language models (VLMs) make this feasible. To advance progress in this area, we introduce the task of navigational sign understanding which parses locations and associated directions from signs. We offer a benchmark for this task, proposing appropriate evaluation metrics and curating a test set capturing signs with varying complexity and design across diverse public spaces, from hospitals to shopping malls to transport hubs. We also provide a baseline approach using VLMs, and demonstrate their promise on navigational sign understanding. Code and dataset are available on Github.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ayush Agrawal, Joel Loo, Nicky Zimmerman, David Hsu. 2025-06-03. Sign Language: Towards Sign Understanding for Robot Autonomy. https://arxiv.org/abs/2506.02556
Cite the original work for its findings. Save a collection to share your selection of sources.