NeFut Logo NeFut
中 Admin Login

[CS.AI] Incident-Arena: Pushing Agents to the Last Nine of Reliability

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #Open Source

AI coding agents are now ubiquitous in engineering workflows across industry and academia. Yet, despite their widespread use in application development, little attention has been paid to their ability to handle production incident response. This emerging area, called agentic site‑reliability‑engineering (SRE), suffers from three major benchmark shortcomings: (1) unrealistic environments, often toy repositories; (2) non‑standard framework implementations; (3) reliance on simple static verifiers.\ We introduce Incident-Arena, a human‑crafted benchmark comprising 20 carefully selected tasks grounded in real‑world deployed open‑source software. Each task deploys a production application to an ephemeral Kubernetes cluster, injects a fault ranging from the configuration layer down to underlying images, and applies a sustained load profile as required. We also present a novel verification approach that goes beyond static checks, employing functional verifiers that keep system‑level metrics stable while ensuring safe repairs.\ Agent trials consume an average of 2.81 M tokens over 41 turns, far exceeding existing benchmarks and demonstrating long‑horizon reasoning. Across the 20 tasks and three application substrates, frontier models score below 64.3%, with failures spanning diagnosis/localization errors, incomplete repairs, and unsafe regressions.\ Review

Original Source: https://arxiv.org/abs/2610.00648

[h] Back to Home