In practice, large language model (LLM) agents are often improved by harness evolution, which refines prompts, tools, and workflows. However, the optimizer’s own failure diagnosis, edit generation, and testing procedures usually stay static. This work investigates whether an optimizer that also evolves its own diagnostic and editing mechanisms can more effectively boost another agent.
Two empirical observations shaped the design:
- In a controlled study, a self‑evolving optimizer fails to improve performance without execution‑based verification; once verification is available, it achieves the best results among all baselines.
- Across five different executors, self‑evolving optimizers automatically construct tools for failure analysis, verification, training audits, and workflow control.
Motivated by these findings, we introduce VERSE (Verified Self‑Evolving optimizer for agent harnesses). VERSE’s workflow consists of:
- Testing draft edits, replaying failures, and perturbing suspected steps before committing;
- Tracking fixes and regressions across iterations;
- Using this feedback to revise both the executor harness and the optimizer’s prompts, skills, tools, hooks, and notes, while keeping the underlying model weights fixed.
Under a shared protocol with disjoint training, validation, and test tasks, VERSE improves all four evaluated harness optimizers on held‑out SWE‑rebench tasks and out‑of‑distribution tasks in five languages. The best validation‑selected harness reaches 42.3% and 37.7% accuracy on validation and test sets respectively, compared to 39.2% and 29.3% for the strongest baselines.
The implementation is publicly available at https://github.com/wzekai/VERSE.
Review