NeFut Logo NeFut
Admin Login

[CS.AI] The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures

Published at: 2026-09-12 22:00 Last updated: 2026-09-15 01:15
#algorithm #AI #Machine Learning

AI agents are increasingly acting through tools and delegated authority, yet general incident repositories rarely capture the mechanisms needed to compare public failures with agent‑security evaluations. We introduce the Agent Incident Registry (AIR), a source‑linked catalog containing \N{} records of agent‑related events disclosed from \Yfirst{} through \Ylast{}. Each entry includes supporting evidence, a stable identifier, and missingness‑aware labels for causal role, disclosure class, mechanism, and outcome. Among the \Nprimary{} generative‑system records where the agent acted, \Rprimary{} involved realized harm, accounting for \Pprimary\% of cases. Realized outcomes cluster in in‑the‑wild and safety‑failure records, while responsible disclosures and research demonstrations are overwhelmingly demonstrative; thus the aggregate share characterizes collection composition rather than deployment risk. After initial curation, a second human reviewer checked all \N{} records and their labels for completeness and correctness. In a deployment‑analogue audit, InjecAgent's \NInjecAgentCases{} cases occupy three of AIR's twelve surfaces and are all attacker‑triggered, whereas AIR contains \Nsafety{} non‑adversary safety failures. AIR supports source‑grounded case retrieval and evaluation‑scope auditing, not failure‑rate or control‑efficacy estimation.

Review

Original Source: https://arxiv.org/abs/2609.11030

[h] Back to Home