NeFut Logo NeFut
中 Admin Login

[CS.AI] Rationale-Guided Policy Optimization: Adaptive Rationale Scaffolding for Learning to Reason

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#Machine Learning #LLM #Artificial Intelligence

On‑policy reinforcement learning has become a central paradigm for enhancing the reasoning ability of large language models, yet its effectiveness is often hampered by reward sparsity. When a model fails to discover correct trajectories on hard problems, the optimizer receives little useful signal and can stall. Existing remedies inject off‑policy demonstrations, expert traces, or model‑generated solutions, but they usually require the auxiliary data to match the RL task format, often relying on rejection sampling from stronger models to obtain suitable trajectories.

We introduce Rationale‑Guided Policy Optimization (RGPO), a framework that adaptively leverages ground‑truth rationales according to the model’s current capability while preserving freedom to explore. Instead of treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model produce improved responses, after which only higher‑reward, model‑generated solutions are fed back into the original unguided setting. This design lets training exploit available rationales without demanding off‑policy data to follow the same format as the RL task.

Across both language‑only and vision‑language reasoning benchmarks, RGPO consistently outperforms RLVR baselines. Ablation studies reveal that adaptive rationale guidance is the primary contributor to these gains. The results suggest that RGPO reduces reward sparsity, stabilizes reinforcement learning, and improves reasoning performance for both text‑only and multimodal models.

Review

Original Source: https://arxiv.org/abs/2610.07342

[h] Back to Home