Large language model (LLM) agents are increasingly merging generation, decision‑making, execution, and self‑evaluation within a single loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications usually remain context for the same model that acts and declares completion, leaving no independent authority boundary. This creates two gaps:
- Understanding‑execution gap: the requirement is understood but not satisfied during execution.
- State‑authority gap: the agent's interpretation or completion claim does not establish the required state.
On SkillsBench, using only agent‑visible prompts, workspace information, and injected skill specifications, we extracted 509 source‑grounded task directions. Across seven models, satisfaction rates ranged from 79.6% to 86.4%, while completion‑claim rates exceeded official evaluator pass rates by 28.7 to 37.9 percentage points.
We therefore separate agent proposals from authoritative state. Agents may plan, act, and request completion, but only admissible evidence from qualified providers can establish a specification‑governed state. SpecHarness operationalizes this principle by compiling visible specifications into source‑linked obligations and governing execution and finalization through versioned obligation state. Verifiable requirements are mediated or validated at runtime, whereas ambiguous or subjective requirements remain advisory. Experiments on guideline‑following and artifact‑generation tasks demonstrate that specifications can serve not merely as behavioral guidance but as authority over compliant execution and completion.
Review