As large language models become common as tool‑augmented legal agents, they introduce agentic hallucinations: errors in tool calls and reasoning cascade into fabricated holdings or mis‑cited authorities. Existing legal benchmarks only measure single‑turn QA outcomes and lack diagnostics for legal agents’ multi‑step trajectories, leaving unanswered how and to what extent an agent hallucinates.
To fill this gap we introduce LexAgentHallu, a benchmark that profiles hallucinations along multi‑step legal agent trajectories. Built via a four‑stage expert‑in‑the‑loop pipeline, it contains 3,414 instances spanning 17 legal categories and 6 task types. Each instance is annotated with a dual‑layer taxonomy: 7 high‑level categories and 27 fine‑grained subclasses, covering both substantive errors and procedural failures of the agent.
We also devise fine‑grained metrics that quantify the severity of each failure and pinpoint its location on the agent’s execution path. Evaluating 18 proprietary and open‑source agents with these metrics reveals a Right‑Answer‑Wrong‑Reason effect and shows that hallucination subclasses tend to cluster rather than scatter, forming distinct profiles tied to the agent framework, legal task, and category. These insights, invisible to outcome‑only evaluation, demonstrate LexAgentHallu’s diagnostic power for legal agentic hallucinations.
Review