Voice agents in production need to juggle multiple requests, background speech, and impatient users. We introduce the VAmoS Energy benchmark, which packs these challenges into 100 calls about utility billing and payment assistance. Each caller issues two to four requests; the agent has sixteen tools backed by a stateful Stripe billing twin and the Apache Fineract loan engine, with account access blocked until verification succeeds. The tasks use public household electricity data and follow Pennsylvania residential billing rules. An LLM‑as‑verifier checks the agent's actions and spoken figures against explicit requirements. In a calibration run it agrees with a code verifier on 99.1% of checks.
Across fourteen voice stacks and three repeats per task, completion rates range from 17.3% to 44.7%. Grok Voice leads, while Gemini 3.8 Live and GPT‑Live achieve similar cost per call. Background television reduces pooled completion from 38.7% to 8.6%. The simulated caller often accepts an incorrect result because it hears the agent’s words but cannot inspect the underlying actions.
These findings demonstrate that voice agents must be evaluated over the entire call, covering what they say, what they change, and how they handle competing speech.
Review