A harness is the code surrounding a language‑model agent that organizes prompts, calls tools, manages context and controls execution. As models become stronger, researchers have begun to let agents improve their own harnesses, a line of work known as self‑evolving harnesses. Existing approaches usually run a separate proposer on a human‑designed harness to modify the solver's harness, and evolve a distinct harness for each benchmark. Real‑world tasks span many domains, so both evolution and evaluation should cover a diverse task set. We propose a framework close to recursive self‑improvement: the same frozen model, on the same harness version, first acts as the solver and then as the proposer, reading the full run records and directly editing the harness that executed them. Each evolution batch draws tasks from five benchmarks in different domains. To measure generalization, training and held‑out tasks are strictly separated, and we additionally evaluate on five out‑of‑distribution benchmarks never used during evolution. We frame the evolution process as deep‑learning training with two stages, multi‑task pretraining and continual training. Starting from a 49‑line seed harness, the harness obtained after the first stage improves the average score by 4.48 points on in‑distribution benchmarks and by 12.64 points on out‑of‑distribution benchmarks, surpassing Codex on the former and matching it on the latter. In the second stage, continued evolution on Claw‑Eval, one of the out‑of‑distribution benchmarks, raises its score from 66.17 to 68.06, exceeding Codex. An in‑depth analysis reveals mechanisms that emerged during evolution, including output truncation, history compaction and independent review.
Review