Voice agents are being deployed in more workflows, and interaction failures can directly affect transactions, access and other critical outcomes. A reproducible and interpretable evaluation method is therefore required. We introduce Inquesto Score (IS), defining reliability as the proportion of calls in a fixed, versioned evaluation population that achieve the caller’s goal without a functional failure or worse. IS does not combine heterogeneous metrics; instead it enumerates explicit failure events, assigns severity levels, and evaluates the deployed voice pipeline directly.\ \ Timing failures such as talk‑over and delayed responses are measured straight from audio, while semantic and state‑dependent failures are judged using scenario predicates, tool traces and a pinned open‑model judge. Diagnostic views of behavior, acoustic robustness, identity handling and speaker groups accompany the score but are not folded into it.\ \ Inquesto Score v0.1 covers 30 scenarios, three acoustic conditions, four speaker groups and runs 306 calls per agent across 13 reference system configurations. Our results show that transcripts alone are insufficient; explicit treatment of deployment conditions and validation of the evaluators are essential for reliable measurement. The protocol, reference implementation and full evaluation records are released.\ \ Review