arXiv · 2604.03701
VidNum: Diagnosing VLM Failure Modes in Video-Grounded Numerical Reasoning
Abstract
Video-grounded numerical reasoning requires Vision-Language Models (VLMs) to identify, track, and combine quantitative evidence across frames, actions, and scene changes. Existing benchmarks provide fragmented coverage: general VideoQA includes counting among broader tasks, while dedicated benchmarks focus on repetition counting, ultra-long-video enumeration, or instructional mathematics. We introduce VidNum, a manually curated and independently verified benchmark containing 1,167 multiple-choice questions. Its three task groups distinguish Direct and Distinct Enumeration, Conditioned and Structured Enumeration, and Compositional Quantitative Reasoning. Question-level annotations further identify the evidence target, counting structure, and required reasoning operation. The best evaluated VLM reaches 59.8% accuracy, compared with 98.2% for human annotators, and no evaluated open-weight model exceeds 45%. Stratified analyses reveal that failures are not uniformly distributed: structured target construction and action-grounded compositional reasoning form recurring bottlenecks across models. Zero-shot chain-of-thought prompting is not a reliable remedy: it recovers some errors but breaks previously correct answers, with effects that vary across models and task structures. VidNum therefore supports diagnostic analysis beyond a single aggregate score.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shaoyang Cui, Lingbei Meng, Yaodi Luo, Peize He. 2026-04-04. VidNum: Diagnosing VLM Failure Modes in Video-Grounded Numerical Reasoning. https://arxiv.org/abs/2604.03701
Cite the original work for its findings. Save a collection to share your selection of sources.