TY - RPRT TI - Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning AU - Xincheng Yao AU - Haobo Fu AU - Weiming Liu AU - Chongyang Zhang PY - 2026 UR - https://arxiv.org/abs/2609.28963 ID - 2609.28963 ER -