arXiv · 2610.10179
Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents
Abstract
Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenyu Huang, Xinyu Hou, Pavlos Vougiouklis, Ruofei Lai, Jeff Z. Pan. 2026-10-07. Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents. https://arxiv.org/abs/2610.10179
Cite the original work for its findings. Save a collection to share your selection of sources.