NeFut Logo NeFut
中 Admin Login

[CS.AI] Reinforcement Learning with Decomposed Subtasks

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#algorithm #AI #Machine Learning

Group Relative Policy Optimization (GRPO) and similar policy‑gradient methods collapse an entire multi‑turn rollout into a single scalar trajectory reward before the policy update. When a task consists of distinct skills and environmental feedback is sparse and delayed, this collapsing loses information: the optimizer must implicitly infer which competency drove the outcome and how to adjust behavior. We argue that the proper primitive is not a better scalar but a decomposition—splitting the trajectory reward along subtasks before it reaches the policy update.

We introduce Reinforcement Learning with Decomposed Subtasks (RLDS), whose core is Subtask‑Decomposed Advantage Estimation (SDAE). SDAE replaces the scalar GRPO advantage by partitioning the trajectory reward into per‑subtask shares according to a fixed taxonomy, computing a group‑relative advantage for each subtask, and distributing per‑token credit by weighting each subtask’s advantage by its importance, concentrating it around the step where a reflection marks that subtask’s execution as consequential.

We evaluate RLDS on four agentic benchmarks: FrozenLake (sparse grid navigation), HotpotQA (multi‑hop QA with a single retrieval tool), ScienceWorld (long‑horizon embodied science), and DeepResearch (long‑form research with four tools and a composite rubric reward). Heterogeneity diagnostics emitted during training reveal where decomposition pays off—gains scale with subtask heterogeneity, achieving the largest improvements on high‑heterogeneity tasks ScienceWorld (+11.5 points, 95% CI [+9.8, +13.3]) and FrozenLake (+9.8 points, [+7.0, +12.8]), while results on HotpotQA and DeepResearch remain within noise, as predicted by the diagnostics. ScienceWorld also becomes more compute‑efficient under RLDS than scalar GRPO (‑10.9% wall‑clock time per step), because long rollouts amortize the fixed reflect‑and‑grade overhead.

Review

Original Source: https://arxiv.org/abs/2609.27035

[h] Back to Home