AI agents combine large language models with external data, file‑modifying tools, API calls, or code execution capabilities. Security breaches arise when adversarial content alters tool usage or when the surrounding software contains flaws such as path traversal or command injection. This work focuses on authorized white‑box pre‑deployment auditing: the auditor can inspect the target repository and run a controlled runtime, yet attacks must still go through the task‑defined attacker interface and be confirmed by an external verifier.
We introduce AgentXploit, a two‑role auditing system that separates repository‑level attack‑path discovery from runtime exploitation. The Analyzer Agent traces attacker‑controlled inputs to sensitive operations and records code‑supported candidate attack paths; the Exploiter Agent turns these paths into concrete attacks and refines them using runtime feedback.
To evaluate the approach we built AgentXploit‑Bench, containing 72 reproducible vulnerabilities across 12 open‑source AI‑agent systems and frameworks. Across three runs, AgentXploit achieved a 59.3% end‑to‑end success rate, compared with 38.4% for Codex. Under a token‑budget‑matched comparison, Codex reached 46.3%. On AgentDojo, where injection points are provided, the Exploiter Agent attained a 79.2% attack success rate versus 52.7% for AgentVigil.
These findings highlight that repository discovery and runtime exploitation constitute distinct challenges in comprehensive agent security auditing.
Review