arXiv · 2608.05485
VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
Abstract
Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric-grounded framework covering five video generation and editing settings. For each input instance, VideoArgus generates an output-blind, sample-specific rubric once and reuses it to evaluate all corresponding candidate videos. The rubric defines concrete criteria, scoring rules, failure modes, and evidence plans, which guide criterion-specific VLM QA and visual tools to produce evidence-grounded criterion scores, rationales, and a diagnostic report. We further construct VideoArgus-Bench, containing 1,026 curated input instances built from 653 high-quality images and 416 high-quality videos, with all benchmark rubrics pre-generated, frozen, and released. On a separate 1,260-video human-alignment set, VideoArgus achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks. Model rankings also remain largely consistent across different rubric-generation and evaluation-VLM backbones. All code and data are released. Visit our project page: https://zzzmyyzeng.github.io/VideoArgus
Explore related subjects
Keep this discovery
Ziyun Zeng, Zixuan Wang, Yongsheng Yu, Hang Hua, Jiebo Luo. 2026-08-06. VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing. https://arxiv.org/abs/2608.05485
Cite the original work for its findings. Save a collection to share your selection of sources.