arXiv · 2604.22829
Lost in Motion: Vision Language Models Fail the Dynamic Gauges Test
Abstract
The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, particularly analog gauges. Vision-Language Models (VLMs) have the potential to provide a general solution for gauge reading and have already shown good performance in instrument recognition. However, performing accurate, real-time gauge readings is a more complex task. This paper evaluates state-of-the-art models, including versions from the GPT-5 (5.4 Thinking and 5.3 Instant) and Gemini 3 (Pro and Flash) families, against a set of simple but realistic dynamic gauge reading scenarios. To facilitate this evaluation, we introduce a novel dataset comprising video sequences of three instruments of different gauge types: circular, linear, and Vernier, under diverse motion and speed profiles. Our findings indicate that the evaluated frontier VLMs, under our specific testing conditions, exhibit a limited ability to interpret needle trajectories and scale semantics, failing to provide the traceability and reliability needed for safety-critical monitoring. The results demonstrate that these models have not yet achieved the performance necessary to be classified as trustworthy synthetic instruments under existing IEEE and ISO standards.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tairan Fu, Francisco Javier Santos-Martín, Javier Conde, Elena Merino-Gómez, Pedro Reviriego. 2026-04-19. Lost in Motion: Vision Language Models Fail the Dynamic Gauges Test. https://arxiv.org/abs/2604.22829
Cite the original work for its findings. Save a collection to share your selection of sources.