AI agents can externalize what they have learned from past tasks into reusable skills such as procedures, checklists, code, or other executable artifacts, and retrieve them when solving new tasks. Self‑evolving skill methods rewrite these skills after each round of practice on training tasks and then apply the updated skill to new tasks of the same type. We ask whether the improvement a skill shows on its training tasks carries over to held‑out test tasks. We evaluate five self‑evolving methods and a one‑shot skill on six benchmarks, keeping the model, the agent, and the train/test split identical for all methods. Among the 21 skills that improve on training tasks, five retain the full improvement on test tasks, thirteen retain part of it, and three retain none. No existing method is best across all benchmarks. Reading the skill content reveals that poor transfer often stems from fixing details that should depend on the task, such as column names or output files, or from turning a fix for a single failure into a rule applied to every task. An LLM judge that reads the skill content ranks finished skills the same way test results do in 86% of pairs, but it predicts the effect of a single edit poorly, so edits still need to be validated by execution. Based on these findings we propose Generalizable Skill Optimization (GSO), which keeps only a guide for writing skills and generates a fresh skill for each task; it achieves the highest scores on all six benchmarks.
Review