NeFut Logo NeFut
中 Admin Login

[CS.AI] A Wrong Turn Does Not Ruin the Journey: Deviation-Guided Skill Self-Evolution for LLM Agents

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#Machine Learning #LLM #Artificial Intelligence

Large language model (LLM) agents increasingly rely on natural‑language skills to tackle complex tool‑use tasks. Such tasks often admit multiple valid solution paths, making it inappropriate to force a failed trajectory to imitate a single fixed successful one. Moreover, a failed trajectory is rarely entirely wrong: the agent may gather useful evidence and make meaningful progress before deviating into an erroneous suffix.

To address this, we introduce SkillPivot, a deviation‑point‑guided framework for skill self‑evolution. SkillPivot identifies the transition from a useful prefix to an erroneous suffix using three signals—execution validity, goal progress, and action diversity. Once the pivot is detected, a stronger teacher model continues from the same prefix, producing a successful alternative under the identical interaction history. By contrasting the student’s failed suffix with the teacher’s successful suffix, SkillPivot generates localized skill updates that preserve already effective guidance while correcting the mistake.

Experiments on ToolQA, LogicBench, and WildClawBench demonstrate that SkillPivot consistently outperforms competing skill‑evolution methods, improves various agent models, and yields compact, transferable skill updates.

Review

Original Source: https://arxiv.org/abs/2609.29154

[h] Back to Home