arXiv · 2610.03286
VenusRL: A Fully Disaggregated Agentic RL System with Priority Scheduling and Scalable Interaction
Abstract
Agentic Reinforcement Learning (RL) trains LLM agents through multi-turn interactions with external tool environments. Its multi-turn nature exposes two system-level bottlenecks unaddressed by existing agentic RL frameworks. First, end-to-end training throughput is constrained by the slowest trajectories to complete, yet optimizing per-GPU utilization alone scatters rollout progress across many groups, delaying the completion of enough groups to unblock the next training step. Second, tool sandboxes are statically over-provisioned by their declared memory ceilings, leaving most physical memory stranded while replicating near-identical state across sandboxes launched from the same prompt. We present VenusRL, a fully disaggregated agentic RL system that addresses both bottlenecks. VenusRL's priority-aware action scheduler uses length-prediction heuristics to identify sample groups whose completion is most likely to unblock the next training step, and pushes them ahead of others across batch admission, KV Cache residency, and cross-worker request orchestration. VenusRL's environment resource manager combines a memory-aware admission threshold with a template-keyed page-sharing pool, packing more sandboxes per node while preserving strict memory isolation via write-protected page table entry aliasing and copy-on-write. Across representative agentic RL workloads, VenusRL achieves up to 4.24x end-to-end training speedup over state-of-the-art baselines and reduces environment cost by up to 89%.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mingjun Zhang, Yucheng Li, Menghao Zhang, Shuyong Zhu, Ping Zhang, Xiaohe Hu, Jun Chen, Zhixin Wang, Xutong Wang, He Liu, Yanmin Jia, Shengrong Zhu, Peng Sun, Mingjie Zhang, Liming Liu, Jinlong Hou, Yuan Cheng, Yujun Zhang. 2026-10-02. VenusRL: A Fully Disaggregated Agentic RL System with Priority Scheduling and Scalable Interaction. https://arxiv.org/abs/2610.03286
Cite the original work for its findings. Save a collection to share your selection of sources.