NeFut Logo NeFut
中 Admin Login

[CS.AI] Learning to Revise Reasoning with Segment-wise On-Policy Distillation

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

On-policy distillation (OPD) improves reasoning of large language models by letting a student receive dense token‑wise supervision from a teacher on its own generated rollouts. The token‑wise supervision, however, does not give a coherent alternative reasoning step, and when the student produces a degenerate prefix the teacher’s supervision remains conditioned on that prefix, potentially reinforcing poor patterns.

To address these issues, we introduce Segment‑wise On‑Policy Distillation (Seg‑OPD). The method intervenes on intermediate reasoning segments: first an uncertainty metric selects student segments; then the teacher rewrites the same context to produce a redraft; finally training combines the original token‑wise OPD loss with a segment‑level contrast loss that encourages the student to prefer the teacher’s redraft over its own segment.

We evaluate on mathematical reasoning and competitive programming tasks. Seg‑OPD consistently yields higher revision success rates than baselines and improves overall reasoning accuracy by about 5.22% on average, with stable gains across diverse models and datasets. The code is released publicly.

Review

Original Source: https://arxiv.org/abs/2610.02703

[h] Back to Home