arXiv · 2510.08936
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
Abstract
Recently, Multi-modal Large Language Models (MLLMs) have demonstrated significant performance across various video understanding tasks. However, their robustness, particularly when faced with manipulated video content, remains largely unexplored. In this paper, we introduce Ro-Bench, the first benchmark for evaluating MLLMs on dynamic out-of-distribution (OOD) counterfactual video test sets. Ro-Bench incorporates high-quality, diverse and temporally relevant video data, by editing Style, Object, Background and their compositions. We evaluated eight recent video MLLMs and found that current models exhibit substantial performance degradation on Ro-Bench when exposed to counterfactual video content. Furthermore, we demonstrate that fine-tuning MLLMs with counterfactual data enhances robustness, achieving a 21.73% performance increase on Ro-Bench and a 12.78% improvement across 20 tasks in the MVBench dataset. These findings underscore the effectiveness of counterfactual data in enhancing the video understanding ability of MLLMs. The code and data will be released shortly.
Explore related subjects
Keep this discovery
Zixi Yang, Jiapeng Li, Muxi Diao, Yinuo Jing, Kongming Liang. 2025-10-10. RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos. https://arxiv.org/abs/2510.08936
Cite the original work for its findings. Save a collection to share your selection of sources.