NeFut Logo NeFut
中 Admin Login

[CS.AI] TISD: On-Policy Self-Distillation with Trajectory Intervention

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#algorithm #AI #Machine Learning

In on‑policy self‑distillation (OPSD), the teacher supplies dense targets but they are evaluated only on trajectories sampled by the student. When a privileged teacher prefers a different action at an already visited prefix, OPSD can give a target for the branching decision yet cannot supervise the successor contexts induced by that action unless the student samples it. This creates a data‑collection bottleneck during training and suggests that teacher‑student disagreement should be used to propose trajectory branches rather than merely to provide local repairs.

We built a diagnostic framework using controlled token interventions and found that a teacher‑preferred token at the point of peak disagreement can improve the student’s continuation success, although its local corrective value is limited. Motivated by this, we introduced a simple branch‑regenerate‑distill pipeline called Trajectory‑Intervention Self‑Distillation (TISD). TISD forces the teacher‑selected branch action, then lets the student generate the suffix, and finally distills the full trajectory under the privileged‑context‑conditioned teacher.

Experiments on coding models show that TISD improves Avg@4 by 1.2 % over SDPO. In scientific domains, under an equal‑step budget Avg@128 rises by 0.8 points and under an equal‑time budget by 0.3 points. These results support teacher‑guided branching as a way to expose useful successor contexts for self‑distillation.

Review

Original Source: https://arxiv.org/abs/2609.30878

[h] Back to Home