NeFut Logo NeFut
中 Admin Login

[CS.AI] Self-Healing Harness for Runtime Oversight of Agent Self-Modification

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

LLM agents can modify their future behavior, raising a fundamental control question: which self‑generated changes should be allowed to persist. We cast this as admission control for self‑modification.

The approach lets an agent propose edits to its operating instructions while an external runtime gate decides on persistence. We built a model‑agnostic self‑healing harness that wraps an otherwise unchanged agent in a Detect‑Notice‑Heal‑Validate loop.

Developers author candidate behavioral rules in an external workspace; these rules receive provisional execution authority during evaluation. Persistent cross‑episode authority is granted only if the rule measurably improves the triggering failure and does not regress protected cases beyond a fixed margin.

Replay supplies matched evidence when available; otherwise forward trials act as a weaker fallback, and a corpus‑level guard re‑tests the accumulated active rule set.

Across 16 matched baseline‑harness pairs spanning AppWorld, Terminal‑Bench, and τ²‑Bench, the gate rejected 383 replay‑decided proposals. Of these, 211 (55%) improved the triggering failure while degrading a previously working case. This shows that locally beneficial self‑modifications often introduce collateral regressions that materially affect gate decisions, providing direct empirical motivation for external admission control.

Task‑completion scores were higher under the harness in all 16 pairs, with two bootstrap intervals excluding zero; repeated‑trial reliability improved in 12 pairs, tied in 4, and never decreased. Because adaptation changes the policy‑inducing context while leaving model weights fixed, admitted changes remain inspectable, reversible, and compatible with closed‑weight models.

Review

Original Source: https://arxiv.org/abs/2609.24130

[h] Back to Home