AI agents are increasingly deployed in real-world settings, interacting with external tools and making sequential decisions under limited human oversight. This creates a pressing demand for reliable, auditable explanations. Conventional Explainable AI (XAI) methods focus on output‑level insights and fall short of delivering the process‑level transparency required for interactive, multi‑step systems. To bridge this gap, we introduce a post‑hoc XAI framework that converts a lengthy agent execution trace into a structured report and a faithful natural‑language explanation grounded in observable behavior. Because it relies solely on execution traces, the framework is applicable across diverse agent architectures, environments, and tasks. Human and automated evaluations on multiple benchmarks and architectures demonstrate that the framework produces high‑quality, trace‑faithful explanations while reliably detecting unsupported claims, unjustified actions, and evidence gaps, outperforming naïve LLM‑generated explanations.
Review