NeFut Logo NeFut
中 Admin Login

[CS.AI] Evaluating Real-Time Voice Agents: From Component Quality to Grounded Outcomes

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#algorithm #AI #Machine Learning

Real‑time voice agents have moved from research prototypes to production, yet the literature is split among speech foundation modelling, turn‑taking psycholinguistics and agentic evaluation, with little cross‑citation. Architecture papers report latency, turn‑taking papers report prediction accuracy, and agentic benchmarks report task success, so a single metric cannot tell whether a deployed agent is truly good. We gathered 38 primary sources into a six‑category, application‑centric taxonomy and derived three evidence‑based claims. First, architecture choice is a deployment constraint rather than a settled verdict: a 2026 enterprise tutorial notes that no fully self‑hosted end‑to‑end system meets production limits, while a chunked cascade independently achieves state‑of‑the‑art duplex behaviour, showing duplex behaviour can be separated from duplex architecture. Second, evaluation has shifted from component quality to grounded outcomes; recent benchmarks verify backend state instead of trusting the agent’s claimed actions. Third, the dyadic assumption is breaking down: multiparty turn‑taking and multi‑speaker reasoning benchmarks reveal that deciding when to stay silent and reasoning about who should receive information are first‑class capabilities that two‑participant frames cannot capture. For each source we state the problem addressed, its mechanism, reported evidence, the search strategy, inclusion criteria, and a verification step that uncovered a misattributed arXiv identifier. We propose the TRG (Timing‑Recovery‑Grounded) reporting standard, characterising an agent by timing, post‑disruption recovery and state‑verified outcome, with an optional fourth axis for multiparty deployments.

Review

Original Source: https://arxiv.org/abs/2609.30798

[h] Back to Home