NeFut Logo NeFut
中 Admin Login

[CS.AI] Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#Machine Learning #LLM #Artificial Intelligence

On‑policy distillation (OPD) improves a student model by mixing student‑generated rollouts with dense token‑level supervision from a teacher, but providing supervision for every rollout incurs heavy teacher computation. We introduce Success‑Referenced On‑Policy Distillation (SR‑OPD), which cuts this cost by selecting which prompts and rollouts receive teacher supervision.

When the student produces both successful and failed rollouts for the same prompt, the successful rollout serves as a natural reference to decide which failures merit teacher feedback. SR‑OPD concentrates on such prompts and prioritizes failed rollouts whose hidden‑state trajectories show sustained divergence from the successful reference, while accounting for the estimated teacher‑input cost.

Across three teacher‑student pairs and six mathematical reasoning benchmarks, SR‑OPD uses only 3.46%‑5.02% of the teacher‑input tokens required by vanilla OPD in a one‑pass setting, yet achieves comparable reasoning performance. In a controlled setting matched to 5% of vanilla OPD’s teacher‑input budget, further experiments confirm two key design choices: (1) focusing supervision on prompts with both successful and failed rollouts, and (2) using successful rollouts to guide failure selection.

These results indicate that a student’s own successful behavior can act as a useful reference for allocating teacher supervision under a fixed teacher‑input budget.

Review

Original Source: https://arxiv.org/abs/2610.02678

[h] Back to Home