NeFut Logo NeFut
中 Admin Login

[CS.AI] Talk2Agent: Benchmarking Voice Interfaces for Text Agents

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Large language model (LLM) computer-use agents are usually evaluated with clean written instructions, yet speech is becoming a dominant way to interact with them. Speech input adds a failure point: transcription errors can alter task‑critical entities, constraints, or targets before the agent starts reasoning, and conventional ASR metrics do not directly measure whether the information needed for successful execution is preserved.

We introduce Talk2Agent, a benchmark that measures how well voice interfaces convey human‑spoken instructions to LLM‑based computer-use agents. Talk2Agent converts tasks from WildClawBench and OSWorld into spoken versions and evaluates a variety of voice interfaces, including dedicated ASR models, audio‑capable LLMs, contextual biasing, and LLM‑based ontology repair.

Because repeatedly running long‑horizon computer-use tasks is costly and stochastic, we propose an execution‑free, task‑conditioned evaluation framework. The framework projects the original task grader onto prompt‑addressable intentions, quantifying how much task‑relevant information survives the voice interface. On WildClawBench, Talk2Agent’s execution‑free native projection offers a practical, execution‑grounded quality measure, correlating with downstream task completion and improving Pearson correlation by 0.246 over WER/CER on 32 hours of real human speech.

Review

Original Source: https://arxiv.org/abs/2609.38867

[h] Back to Home