Searcharxiv⌕ Search

arXiv · 2609.28963

Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

Abstract

Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xincheng Yao, Haobo Fu, Weiming Liu, Chongyang Zhang. 2026-09-24. Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning. https://arxiv.org/abs/2609.28963

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Categorical Approach to Conflict Resolution:A Corrected Correspondence between Category Theory and the Graph Model for Conflict Resolution

This note is a substantially revised version of the author's earlier preprint (version~1, 2023), which proposed a ``Categorical Graph Model for Conflict Resolution'' (C-GMCR). Several of the central definitions and claims of version~1 are incorrect, and this version replaces them. We first observe that the states and one-step moves of a graph model do not form a category; the correct construction is the free category on a quiver whose arrows are moves labelled by the decision maker (DM) who controls them. Controller labels are functorial: they define a functor into the free monoid on the set of DMs, and the legal move sequences underlying coalition reachability are exactly the paths whose label words contain no immediate repetition. Preferences, in contrast, are not functorial. We show that a map representing a DM's preference in a preordered set exists only if the preference is transitive, and that even then it extends to a functor on the path category if and only if every move, regardless of its controller, is weakly improving for the focal DM; in that case general metarationality, symmetric metarationality and sequential stability all collapse to Nash stability for the focal DM. Preferences therefore enter the categorical picture as a selection of a subquiver of improvements, not as a functor. We restate the four standard stability concepts in path language, work them out on the Prisoner's Dilemma, and take a first step towards comparing graph models via induced embeddings: Nash stability is reflected by induced embeddings, whereas general metarationality, symmetric metarationality and sequential stability are neither preserved nor reflected. A section lists each correction to version~1 explicitly.

cs.AI↗

From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark

Recent reasoning-oriented LLMs have demonstrated strong performance on challenging tasks such as mathematics and science examinations. However, core cognitive faculties of human intelligence, such as abstract reasoning and generalization, remain underexplored. To address this, we evaluate recent reasoning-oriented LLMs on the Abstraction and Reasoning Corpus (ARC) benchmark, which explicitly demands both faculties. We formulate ARC as a program synthesis task and propose nine candidate solvers. Experimental results show that repeated-sampling planning-aided code generation (RSPC) achieves the highest test accuracy and demonstrates consistent generalization across most LLMs. To further improve performance, we introduce an ARC solver, Knowledge Augmentation for Abstract Reasoning (KAAR), which encodes core knowledge priors within an ontology that classifies priors into three hierarchical levels based on their dependencies. KAAR progressively expands LLM reasoning capacity by gradually augmenting priors at each level, and invokes RSPC to generate candidate solutions after each augmentation stage. This stage-wise reasoning reduces interference from irrelevant priors and improves LLM performance. Empirical results show that KAAR maintains strong generalization and consistently outperforms non-augmented RSPC across all evaluated LLMs, achieving around 5% absolute gains and up to 64.52% relative improvement. Despite these achievements, ARC remains a challenging benchmark for reasoning-oriented LLMs, highlighting future avenues of progress in LLMs. Our code is available at https://github.com/you68681/kaar.

cs.AI↗

GLOVE: Global Verifier for LLM Memory-Environment Realignment

Most existing memory-enhanced Large Language Model (LLM) approaches implicitly assume that memory validity can be established either through external evaluators that provide task-specific success signals or through internal model cognition, such as reflection, for editing memory entries. However, these assumptions often break down in practical environments with dynamic drifts. We propose the Global Verifier (GLOVE), a framework that introduces a new design dimension for LLM memory systems by establishing a relative notion of truth. Through active probing to detect inconsistencies between retrieved memories and fresh observations, GLOVE enables memory-environment realignment by verifying and updating memory without access to ground-truth supervision or strong reliance on model introspection. We evaluate GLOVE on diverse benchmarks spanning web navigation, planning, and control, augmented with controlled environmental drifts that introduce non-stationarity beyond the original benchmark settings. Our results show that GLOVE substantially improves agent success rates, suggesting a robust pathway to cognitive agents capable of self-evolving.

cs.AI↗