NeFut Logo NeFut
Admin Login

[CS.AI] Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

Published at: 2026-09-11 22:00 Last updated: 2026-09-12 06:35
#AI #Machine Learning #LLM

Agentic systems are rapidly moving into production, where they ingest untrusted inputs, invoke tools with real permissions, and act autonomously, expanding the attack surface beyond chat‑only models. Existing evaluations are mostly single‑turn and miss multi‑step agent vulnerabilities. We introduce a black‑box risk‑aware evaluation framework that requires only basic system descriptions. The framework comprises a seven‑domain taxonomy that maps observable behaviors to risk categories, a fully automated SAGE‑RT red‑team generator that creates 120 adversarial scenarios per domain, and human‑validated assessment using LLM judges. Empirical tests on two agent architectures (CrewAI and AutoGen) with four base models reveal alarming numbers: average governance risk of 56.25%, privacy risk of 65% in multi‑agent setups, and behavior vulnerabilities reaching up to 85%. This black‑box approach identifies critical architectural flaws without privileged access, offering a scalable path toward safer agent deployments.

Review

Original Source: https://arxiv.org/abs/2609.09647

[h] Back to Home