We introduce Looped Self-Distillation, a self‑evolution framework for code generation where a model repeatedly produces and learns from its own raw outputs under a fixed information budget, without external evaluation or test‑based sample selection. A key observation is that correctness can improve while the breadth of correct implementations shrinks. To address this, we propose SPECTRUM, which re‑estimates loss‑sensitive key/value geometry from a fixed reference anchor at each round and converts it into a full‑rank proximal spectral modulation. All generated completions train a single student model, whose subsequent inference requires no further intervention. In five rounds of experiments on MBPP, SPECTRUM retains 89.9% of the initial model's 64‑sample correct AST richness, compared with 66.4% for vanilla self‑distillation and 65.5% for a subspace‑projection control. The advantage persists when matching correct‑sample counts. Without additional training or recalibration, the resulting student also achieves higher matched‑correct richness than vanilla SD on HumanEval+ and APPS Intro, demonstrating transfer of the diversity benefit. These results establish correct‑solution retention as a complementary objective of recursive self‑improvement and show that generation‑time intervention can improve the solution repertoire retained by subsequent students.
Review