Tool-using language‑model agents frequently mutate schedulers, data pipelines, object stores and access‑control systems. Between a read and its commit the external state may be changed by other actors, but not every change makes the commit unsafe. We distinguish invalidating races that break a declared safety predicate from predicate‑preserving and irrelevant races, and ask how precisely runtime guards can separate them.
Our deterministic simulator separates visible from authoritative state and injects five non‑atomic failure mechanisms across sixteen infrastructure tasks in four domains. Frozen agent proposals are replayed counterfactually under every controller without an LLM judge.
Experiments run 3,456 trajectories on three locally hosted quantized models (Qwen3‑4B, Phi‑4‑mini, Gemma4‑8B) and evaluate three commit‑time guard granularities—global epoch, read‑set version and semantic commit predicate—combined with multi‑level verification and model‑side gates.
All three guards eliminate unsafe commits. Freshness‑based guards unnecessarily block 92‑95% of benign races, forfeiting up to 43% of safe task completions, while the complete predicate guard blocks none. Precision is contract‑dependent: deleting a single declared clause turns its entire fault family into unsafe commits (up to 7.9%). Model‑side signals do not substitute: verbal confidence is miscalibrated (ECE≈0.37), action agreement matches a random gate, a cautionary prompt leaves the direct unsafe rate essentially unchanged, and after a freshness‑guard block agents re‑commit unsafely from refreshed but still‑incomplete reads. Under degraded telemetry a hidden concurrent mutation remains observationally clean, bounding every selective policy. Hence precise runtime enforcement requires semantic contracts rather than freshness heuristics or model self‑assessment.
Review