NeFut Logo NeFut
中 Admin Login

[CS.AI] Gains and Collapse in On-Policy Distillation: A Reinforcement Learning Perspective

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

On‑policy Distillation (OPD) has become a cornerstone for post‑training large language models. While OPD often yields performance gains, it can also collapse into overly long and repetitive generations, and the cause of these divergent outcomes has been unclear. We interpret OPD through a reinforcement‑learning lens: the teacher model implicitly rewards student behaviors, even those it rarely produces itself. Our experiments show that OPD improves performance without expanding the student’s capability set; when the implicit reward model is trustworthy, correct responses become easier to sample. Conversely, if the preference encoded in the implicit reward misaligns with true quality, reward hacking occurs—the implicit reward amplifies over‑long, repetitive roll‑outs although the teacher seldom generates such text. Guided by this diagnosis, we propose two straightforward mitigations: masking unhealthy responses during training and initializing the student with supervised‑fine‑tuning (SFT). Both strategies substantially reduce collapse in our evaluations. In short, OPD magnifies student behaviors favored by the teacher’s implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student roll‑outs. The code is released at https://github.com/HancCui/opd_hacking

Review

Original Source: https://arxiv.org/abs/2610.03185

[h] Back to Home