NeFut Logo NeFut
Admin Login

[CS.AI] τ-Elicitation: Benchmarking Multi-turn Entity Extraction in Voice Agents

Published at: 2026-09-15 22:00 Last updated: 2026-09-16 00:22
#algorithm #AI #Machine Learning

Voice agents must capture names, addresses, identifiers, dates, and times with exact precision, yet end‑to‑end benchmarks often hide where capture fails. To address this, we introduce $\tau$-Elicitation, a benchmark comprising 200 tasks across 10 entity types, with controlled difficulty, caller realisms, and three environments. A matched text‑only agent solves all tasks, whereas four voice configurations achieve robust exact success rates between 0.14 and 0.41. Agents increase verification for hard or unfamiliar entities and sometimes for incorrect captures, but not for the weakest caller voice; only 24% to 37% of verified errors are repaired. Adding a scaffold that enforces spelling, read‑back, correction, and confirmation raises robust $Pass^3$ by 14 to 31 points, at the cost of 21 to 28 seconds per call. Realisms such as spelling variations and restarts do not noticeably affect exact success, while mispronunciation raises repair effort. These findings identify strategy selection and successful recovery as the central bottlenecks for exact spoken entity collection.

Review: The benchmark offers a fine‑grained evaluation framework for voice agents, highlighting the critical role of verification and error‑recovery strategies in real‑world dialogues.

Original Source: https://arxiv.org/abs/2609.13602

[h] Back to Home