Existing benchmarks for computer‑use agents mainly test retrieval, overlooking the full workflow an assistant must handle. A useful assistant needs to gather information across complex, multi‑step processes, synthesize it into artifacts such as documents, presentations, or spreadsheets, and interact with program interfaces to produce a coherent final output. This demands reasoning, task decomposition, and visual‑spatial understanding.
To study agents on such workflows we introduce KNOWS, a benchmark of open‑ended, complex, browser‑based tasks that all terminate in a concrete artifact. Task creation follows a design rubric and a protocol that guarantees each task meets the required criteria. Every task is paired with an evaluator program that blends deterministic checks with large language model (LLM) judgments, balancing richness, reliability, and automation in agent evaluation.
We evaluated state‑of‑the‑art computer‑use agents and browser‑based harnesses. Agents achieved moderate scores on partial‑success metrics, yet the best performer fully succeeded on fewer than 3% of our complex, long‑horizon tasks. Even when agents completed over 50% of the other evaluation steps, failures on visual steps rendered the resulting artifacts unusable.
These results expose the limitations of current agents acting as end‑to‑end assistants and call for progress in tool use, visual understanding, and long‑horizon reasoning.
Review