arXiv · 2510.24126
Reinforcement Learning for Long-Horizon Multi-Turn Search Agents
Abstract
Large Language Model (LLM) agents can leverage multiple turns and tools to solve complex tasks, with prompt-based approaches achieving strong performance. This work demonstrates that Reinforcement Learning (RL) can push capabilities significantly further by learning from experience. Through experiments on a legal document search benchmark, we show that our RL-trained 14 Billion parameter model outperforms frontier class models (85% vs 78% accuracy). In addition, we explore turn-restricted regimes, during training and at test-time, that show these agents achieve better results if allowed to operate over longer multi-turn horizons.
Explore related subjects
Keep this discovery
Vivek Kalyan, Martin Andrews. 2025-10-28. Reinforcement Learning for Long-Horizon Multi-Turn Search Agents. https://arxiv.org/abs/2510.24126
Cite the original work for its findings. Save a collection to share your selection of sources.