Legal research consumes a large portion of a lawyer's workflow: identifying controlling authority, confirming its validity, reconciling statutes and cases, and producing a grounded answer. Language‑model agents fit this retrieval‑heavy process, and even partial automation would be valuable. However, reliability is a prerequisite—any missing authority, stale citation, or wrong conclusion can render an answer unusable.
We introduce the Legal Research Bench (LRB), a collection of 413 open‑ended U.S. legal research questions authored by experts. Each question includes a gold answer, supporting authorities, and a binary grading rubric. Thirteen frontier models are evaluated within a harness that integrates web search, case‑law search, page parsing, and retrieval tools.
Scoring follows an all‑pass criterion: a response is deemed correct only if every required condition is met and the cited authorities are verified. An LLM judge is validated against experienced attorneys to ensure benchmark scores align with professional judgment.
Results show that current agents are far from reliable. The strongest model, Claude Opus 4.8, achieves full correctness on just 42.9% of questions. Performance varies widely across legal domains, with lower success on tasks that require reconciling conflicting authorities. More dialogue turns, tool calls, or higher inference cost do not reliably predict higher accuracy.
Review