A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone is insufficient for intervention because the replacement itself may introduce the failure the verification aims to prevent. TwinCheck introduces an inference‑time verification policy that only considers replacement when the trace satisfies an evidence condition tied to a trace‑local failure hypothesis. It constructs a trace‑grounded counterfactual alternative, called a negative twin, and replaces the agent's proposal only if the twin passes structural checks and a pairwise verifier prefers it in both candidate orders.
For paired evaluation, Exact Replay holds the agent's parsed responses and actions fixed until the first accepted replacement, separating intervention effects from resampling. In the primary analysis of 159 multi‑turn BFCL V4 tasks with complete Exact‑Replay pairs, the full policy raises GPT‑5.6 Sol's task success from 45.3% to 58.5% (95% task‑bootstrap CI [8.2, 18.8]), with no observed regressions from success to failure.
These findings recast execution‑boundary repair as a constrained comparison, making the counterfactual action itself the object of verification.
Review