Deterministic benchmark scores indicate that an agent earned credit, but they do not reveal whether the credit was truly earned, reported honestly, or would persist on a second run. We introduce a Trust Layer—an additive post‑hoc framework that, alongside each recorded score, reports whether the score should be believed. The layer checks four properties: (1) the result conforms to the benchmark’s own grading logic; (2) a passing answer was obtained through traceable computation; (3) the agent’s completion claim matches what actually happened; (4) the result remains stable under repeated execution. The first three checks rely only on saved artifacts, while the fourth re‑runs the agent. Model judgments label evidence under majority voting; all verdicts follow deterministic rules and never alter the recorded score.
We applied the Trust Layer to five agent configurations on 108 tasks from Agents' Last Exam. Every model exhibited passing runs without traceable computation (rates varied by an order of magnitude), confirmed false completion claims, and unstable outcomes—18%‑46% of tasks changed score bands across five runs. Only 22.6% of recorded passes cleared all four checks (95% CI 15.0‑32.6, n=84).
The findings highlight that measuring what an agent can do and verifying that it actually did it are distinct problems, and current benchmarks address only the former.
Review