Language model agents are usually wrapped with a harness that manages context, tools, and feedback. When a large‑scale teacher agent is distilled into a smaller student, the harness stays unchanged, so the student only needs to acquire abilities that go beyond what the harness provides, such as correctly interpreting harness information. Standard distillation simply imitates the teacher’s full output and treats the harness as ordinary input, which prevents the student from learning the teacher’s extra contribution. To address this, we introduce Harness‑Aware Distillation (HAD), which isolates the teacher’s value beyond the harness. HAD augments on‑policy distillation with two components:
- Action preference: contrasts the teacher’s actions with and without harness information, scoring the preference after the student’s own reasoning;
- Validity check: discards preference pairs whose preferred action contradicts the harness records.
The contrast supplies information that pure imitation cannot provide, and HAD requires no task rewards, success labels, or future information. Experiments on several long‑horizon agent benchmarks and across different model sizes show that, with the same fixed harness, HAD consistently outperforms on‑policy distillation baselines. Further analysis reveals that HAD enters fewer unproductive loops and recovers from errors more often than baselines, suggesting it retains learnable feedback in its weights while still reading state information from the harness.
Review