arXiv · 2507.15824
PhysVidBench: Language-Grounded Evaluation of Physical Commonsense in Text-to-Video Models
Abstract
Text-to-video (T2V) models now produce striking visuals, yet they routinely violate everyday physics; objects float, tools are misused, and causal sequences break down. Existing benchmarks mostly probe isolated physical laws and rely on vision-language models to score videos directly, which entangles perception and reasoning in one judgment and correlates poorly with humans. We introduce PhysVidBench, a benchmark of 383 base prompts, expanded to 766 prompts with enriched variants and 4,123 manually reviewed QA items, together with a language-grounded evaluation framework that takes a different route: instead of asking a VLM "Is this video physically correct?", we caption the video, then ask a language model to answer prompt-derived yes/no physics questions using only the captions. This split between seeing and reasoning makes each score traceable to the captions that justify it and aligns more closely with human judgment than direct VLM scoring (Pearson r=0.45-0.69, compared with 0.21-0.47 for the strongest direct VLM evaluator). To test the framework, we carefully curate a set of human-validated prompts spanning seven physical dimensions, with a focus on tool use and affordances - areas largely absent from prior physics-focused benchmarks. Across 12 open and proprietary T2V systems, the best model reaches only 36.2% average accuracy, and no model consistently handles everyday physical reasoning. The same pipeline can also guide iterative error-guided prompt refinement, improving CogVideoX-2B from 21.6 to 32.7 and CogVideoX-5B from 17.8 to 29.7 without retraining.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Enes Sanli, Baris Sarper Tezcan, Erkut Erdem, Aykut Erdem. 2025-07-21. PhysVidBench: Language-Grounded Evaluation of Physical Commonsense in Text-to-Video Models. https://arxiv.org/abs/2507.15824
Cite the original work for its findings. Save a collection to share your selection of sources.