Tool‑calling agents interleave structured tool invocation tokens with user‑facing natural language summaries, creating heterogeneous outputs that expose a structural failure mode in standard on‑policy reinforcement learning. Algorithms such as GRPO indiscriminately broadcast a single trajectory‑level scalar advantage to every token, so gradient noise from summary generation contaminates tool‑decision tokens, leading to cross‑segment credit misattribution and brittle optimization.
This work introduces SLCA‑GRPO, a framework that incorporates Segment‑Locked Credit Assignment (SLCA). We first build the Schema‑Guided LLM Simulator (SGLS) as a scalable training substrate, avoiding costly real‑API calls while keeping training stable. SLCA decouples advantage estimation at the structural segment level within a single batch of rollouts, eliminating the need for additional rollouts from intermediate states.
Hierarchical Rewards (HierR) route execution advantages exclusively to tool tokens and preference advantages exclusively to summary tokens, removing the dominant cross‑segment credit misattribution channel in each policy update. On a 7B backbone, SLCA‑GRPO converges faster and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in‑domain evaluation, +1.36 pp on the Berkeley Function‑Calling Leaderboard (BFCL), and +9.15 pp on $\tau^2$‑Bench under identical training budgets, achieving higher accuracy with reduced tool redundancy and cost.
Review