NeFut Logo NeFut
中 Admin Login

[CS.AI] Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #Graph

Group-based reinforcement learning methods such as GRPO and its variants provide reliable group‑normalized advantage estimates at the response level, but they become systematically biased at the step level because coarse trajectory advantages cannot accurately capture the contribution of individual steps—failed trajectories may still contain valuable actions.

Revisiting the basic RL definition reveals that GRPO’s success on single‑turn tasks stems from an advantage estimation that follows the principle: the mean reward of multiple actions sampled from the same state serves as a credible state‑value estimate. Extending this faithful estimation to every intermediate state would require sampling many actions per state, which is prohibitively costly.

We therefore propose the Graph‑based Faithful sTep‑level credit‑assignment framework (GRAFT). All rollout trajectories are grafted into a trajectory graph; Bellman iteration on the graph recovers node state‑values, and the difference between adjacent node values assigns credit to each edge. Theoretically, the resulting step‑level advantage adheres strictly to the fundamental RL advantage definition.

To further mitigate bias in state‑value estimation, we extend GAE to the trajectory graph, creating Graph GAE, which reduces the impact of estimation errors. Experiments on a suite of multi‑turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. The code will be released at https://github.com/xcyao00/GRAFT.

Review

Original Source: https://arxiv.org/abs/2609.28963

[h] Back to Home