Supervised fine‑tuning (SFT) on offline agent trajectories is the standard way to train specialized tool‑using agents, yet forcing token‑by‑token imitation can harm the base model’s general reasoning, tool calling, and code generation abilities. This work investigates how to better balance the trade‑off between acquiring new capabilities and preserving existing ones during agent trace SFT. Comparing several baselines, standard SFT improves the target benchmark but lowers scores on several non‑target benchmarks; merely constraining distributional drift with a KL penalty or limiting update magnitude does not stop this regression. Inspired by recent token‑wise adaptive learning objectives, we propose Privilege‑Guided SFT (PG‑SFT), which uses turn‑level information gain from agent trajectories as a signal to adjust supervision strength. PG‑SFT achieves a more favorable trade‑off across evaluated benchmarks, substantially reducing distributional drift and broad capability degradation while incurring only a slight drop in target‑task performance. The findings suggest that balancing acquisition‑retention depends not only on anchoring the model to its base behavior but also on where and how strongly supervision should diverge from that behavior.
Review