Black-box On-Policy Distillation (OPD) aims to improve a student model using its own generations when the teacher can only return sampled responses without token probabilities. Adversarial distillation introduces a discriminator that scores prompt‑matched teacher and student replies, using the score as the policy reward. Sampling discriminator negatives from the latest student at each step ties the reward to a distribution that shifts after every policy update, creating a moving‑target problem.
The proposed solution, persistent‑negative adversarial distillation, adopts a live‑pool approach: each discriminator batch replaces a fraction of fresh negatives with historical teacher‑student matches. This keeps discriminator compute constant while training the discriminator on historical pairs, and GRPO remains on‑policy with fresh student outputs.
Analysis identifies the Bayes‑optimal reward as the log‑density ratio between teacher and negative distributions. Under explicit assumptions, persistent negatives anchor the discriminator’s negative distribution, reducing reward‑estimation mean‑squared error compared to training with only fresh negatives.
Empirical evaluation spans two student families, three judges, and four judged‑chat benchmarks. With matched discriminator compute, persistent‑negative adversarial distillation consistently outperforms existing methods and yields smoother discriminator trajectories, with fewer below‑chance dips. These results highlight the negative distribution as a crucial design axis in black‑box on‑policy distillation.
Review