Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes.
This paper introduces a property-level reconstructability metric over eight decision-property classes as an evaluation-validity metric, ensuring that captured evidence can reconstruct the decision a claim depends on.
It specifies a counterfactual-replay intervention protocol, implements its replayability-precondition probe, and defines a claim-evidence overclaim gap.
On public and bundled traces, without new model runs, twelve-field sufficiency spans 0.458-0.833 across four inputs sharing a surface reading; replay preconditions are unmet in every scored trace.
In a synthetic release-gate pair, the sufficiency gate blocks the raw variant (0.542) and passes the instrumented (0.667).
Safety-evaluation claims should travel with their reconstructability vector; a reproducibility package regenerates every reported number.
Blogger's Review: The reconstructability metric proposed in this paper holds significant practical value, particularly in evaluating agent safety, where ensuring the reliability and reproducibility of assessment results is crucial.
The introduction of Evidence Sufficiency Cards and counterfactual replay protocols enhances our understanding and validation of evaluation outcomes, advancing the standards of safety assessment.