Language agents must adapt their planned actions when evidence changes, permissions are revoked, or a stop instruction arrives. The desirable response is selective: pause the affected actions, keep the unaffected work intact, and resume only after the necessary repairs.
We introduce NAQD‑Env, a synthetic environment that evaluates such decisions against a deterministic reference policy using explicit evidence, authorization, and constraint dependencies. The environment defines eleven dependency families, allowing evaluation on development structures, held‑out structures, and held‑out combinations of structures.
Metrics distinguish attempted violations from those permitted by a simulated execution gate and jointly report policy agreement, task value, withdrawal, resumption, and event‑reporting statistics. We test three open‑weight instruction‑tuned models (from two model families) under three prompt conditions on 350 frozen scenarios, yielding 3,150 model‑prompt episodes before gate replay.
Across conditions, withdrawal recall never exceeds 0.06, no valid resumption occurs at eligible moments, and only a single episode matches the full reference policy. Under the NAQD prompt, Qwen2.5‑7B produces fewer unsafe‑attempt episodes than Qwen2.5‑3B and Llama‑3.1‑8B, but it completes less useful work and preserves unaffected actions less accurately. Supervised fine‑tuning experiments raise Qwen2.5‑3B’s decision accuracy from 0.45‑0.54 to 0.83‑0.92; however, diagnostics reveal inappropriate withdrawals after curriculum omissions and a loss of event‑reporting capability.
These results motivate treating selective withdrawal as a distinct component of agent reliability. The setting measures policy application with trusted structured inputs and does not imply real‑world containment or source‑verification capabilities.
Review