Large Language Model (LLM)‑based agents evolve themselves by editing a persistent skill document that encodes workflow, tool‑use rules, and decision logic. The evolution loop consists of two steps: an optimizer proposes a candidate edit, and a gate decides whether to accept it. Prior work has focused on the optimizer, while the gate typically follows a naïve rule—accept any edit that improves the aggregate validation score. We identify two fundamental flaws in this rule:
- Permanent regressions – an edit may raise the average score but break items the skill already solves, leading to overall degradation.
- Optimizer’s Curse – on a finite, noisy validation set the best observed score is upward‑biased, causing the gate to over‑accept.
To address these issues we introduce SAGE (Statistical Acceptance Gate for self‑evolving agents). SAGE contributes two key ideas:
- Per‑item paired comparison – the current skill and the edited skill are evaluated on identical validation items, exposing regressions hidden by aggregate metrics and penalizing them asymmetrically.
- One‑sided paired test – an edit is committed only when its wins are statistically reliable against its losses; otherwise the gate abstains. At a boundary setting SAGE exactly recovers the baseline gate, making it a conservative refinement.
We evaluate SAGE on five benchmarks using four backbone LLMs under an equal‑budget protocol. In 19 of 20 settings SAGE reduces the regression rate, matching the baseline in the remaining case. Notable drops include LiveMath (36.5% → 0%) and OfficeQA with DeepSeek‑V4 (42.8% → 0%). Moreover, SAGE achieves the highest final score in all 20 settings, e.g., raising LiveMath from 34.15 to 48.78.
Review: By leveraging statistical paired testing, SAGE curtails the risk of indiscriminately accepting edits while still capturing genuine performance gains, offering a robust gating mechanism for self‑evolving agents.