NeFut Logo NeFut
中 Admin Login

[CS.AI] Dependency-Aware Reward Shaping for Agentic Reinforcement Learning

Published at: 2026-10-03 22:00 Last updated: 2026-10-06 12:11
#AI #Graph #LLM

When training large language models with reinforcement learning, terminal rewards alone provide little guidance about which steps matter. Common step‑credit methods ignore that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure signal, every step in a failed episode receives zero future reward, even if it made progress.

We introduce Dependency‑Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step‑level credit over the resulting dependency graph. An annotator marks for each step which predicates are verified, invalidated, or repaired. Verified predicates are discounted according to their graph distance from the nearest broken prerequisite, while independent predicates are unaffected. Repairs update these weights based on any remaining errors; invalidated predicates must be re‑verified to regain credit. A fixed potential function converts these annotations into signed per‑step rewards.

A common reward and annotation interface lets DARS integrate with a variety of reasoning and agentic training methods such as GiGPO and ARPO/AEPO without altering rollout strategies or optimizers. Across five task families and models ranging from 1.5B to 8B parameters, DARS improves success rates by up to 10 points over GiGPO under the same budget and ALFWorld harness, raises WebShop scores and Search‑R1 QA accuracy, complements AEPO’s entropy‑based training on AIME24/25 with a Python interpreter, and surpasses OmniOPD in controlled tool‑free reasoning comparisons at 1.7B and 4B scales. Ablations show that step‑level credit, dependency attenuation, and graph topology each contribute. On ALFWorld, a distilled 8B annotator matches the API annotator, enabling DARS to run efficiently without a frontier judge. Code is released publicly.

Review: DARS’s explicit modeling of task dependencies provides finer‑grained credit assignment, leading to notable gains in success rates and generalization for large‑scale language agents.

Original Source: https://arxiv.org/abs/2610.01207

[h] Back to Home