arXiv · 2609.13258
Interpretable Temporal Video Reasoning with EventGraph and EventField
Abstract
We present a structured temporal video reasoning pipeline built around a discrete EventGraph, a continuous EventField, and a human-readable EventGlyph view. On a calibrated EPIC-KITCHENS subset of 10 videos and 50 temporal reasoning questions, EventField+Glyph achieves 0.98 overall accuracy, which is higher than the caption baseline by +0.40 (paired p = 1.1 \times 10^{-5}) and direct VLM-only QA by +0.20 (p = 0.0063) on this subset. We further evaluate annotation-source variations, including manual, heuristic, and heuristic+Gemini pipelines, and find that the best structured method stays above the caption baseline across settings. We also include cross-video pair benchmarking and an appendix gallery of glyph outputs for all studied videos. Overall, the results indicate that structured temporal representations can support both performance and inspectability by preserving symbolic structure, capturing temporal continuity, and providing human-readable diagnostics for video reasoning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Durgendra Narayan Singh. 2026-09-06. Interpretable Temporal Video Reasoning with EventGraph and EventField. https://arxiv.org/abs/2609.13258
Cite the original work for its findings. Save a collection to share your selection of sources.