arXiv · 2609.34547
ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models
Abstract
Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gueter Josmy Faure, Min-Hung Chen, Hao Ping Wang, Timothée Lardy, Hung-Ting Su, Winston H. Hsu. 2026-09-28. ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models. https://arxiv.org/abs/2609.34547
Cite the original work for its findings. Save a collection to share your selection of sources.