arXiv · 2610.04911
VideoResearchAgent: Grounded Task Synthesis and Sim-to-Real RL for Open-Web Video Research
Abstract
Existing deep research agents are designed primarily for text- and image-based web sources, while video reasoning systems typically assume that relevant videos are provided in advance. We study open-web video research, where an agent must autonomously discover relevant videos, navigate their temporal content, and ground answers in visual evidence. Training such agents at scale is challenging as live video interaction is slow and unreliable, whereas fixed local simulation can induce retrieval-specific shortcuts that fail to transfer to the open web. We introduce VideoResearchAgent, a scalable training framework to address these challenges. First, we introduce controllable task synthesis pipeline to synthesize multi-hop research tasks from timestamped visual evidence while filtering text-only shortcuts. Second, we build a field-aligned local video simulator that preserves deployment-facing search and watch interactions while accelerating video search by a factor of 34.5-64.6. Third, we introduce Retrieval-Domain-Randomized GRPO (RDR-GRPO), which diversifies candidate rankings, distractors, metadata, and result structure during training to reduce overfitting to simulated retrieval. On Video-BrowseComp, the VideoResearchAgent trained using Qwen3.5-4B achieves 40.48% accuracy, comparable to Gemini-3-Flash-Preview, while reducing cumulative API-token consumption by 74.9% relative to the untrained model. Together, these results establish an accurate and efficient training recipe for open-web video research.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuhang Zhou, Fei Li, Yuxi Wu, Bin Zhu, Jingjing Chen. 2026-10-04. VideoResearchAgent: Grounded Task Synthesis and Sim-to-Real RL for Open-Web Video Research. https://arxiv.org/abs/2610.04911
Cite the original work for its findings. Save a collection to share your selection of sources.