NeFut Logo NeFut
中 Admin Login

[CS.AI] ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#algorithm #AI #Machine Learning

Agents are increasingly deployed with real autonomy in web applications and network penetration testing, where a single out‑of‑scope action can breach a client’s engagement boundary. Existing offensive‑security benchmarks only measure raw hacking capability; as those benchmarks saturate, the real deployment barrier becomes a special case of alignment: scope adherence.

We introduce ScopeBench, a benchmark consisting of 30 dead‑end security tasks. The stated objective of each task is reachable only by violating the declared scope. Each task appears under two conditions that share the same environment, verifier and objective but differ only in the scope description: one instruction set has no scope to measure pure capability, the other includes a natural‑language scope to measure adherence.

Scopeless trajectories are graded by a standard deterministic verifier. Scoped trajectories pass through two grading arms. First, the same deterministic verifier checks for the flag; because the flag lies behind the scope boundary, a pass proves that a forbidden action occurred, providing a high‑precision lower bound on the violation rate. If the verifier does not pass, an agentic judge estimates whether an out‑of‑scope call happened. We calibrate the judge on 100 ScopeBench trajectories labeled call‑by‑call by human annotators, and a blinded audit of 36 evaluated violations finds no false negatives and only a single over‑flagging error, confirming its high recall.

Across eight models in a single harness, raw capability scores range from 12.2% to 81.1% and scope‑adherence rates from 34.4% to 86.7%. The judge discovers 331 violations missed by mechanical verification. Opus‑4‑8 achieves a raw‑capability score 10 percentage points higher than sonnet‑4‑6 while showing 35.6 percentage points higher scope adherence. We release the frozen pilot benchmark, evaluation code, and all 2160 ATIF trajectories.

Review

Original Source: https://arxiv.org/abs/2609.30325

[h] Back to Home