JusticeAxis is a multimodal benchmark for real‑world criminal cases, containing 256 cases from 18 countries with audio, image and text evidence. For each case, three lawyer‑written judgments are provided: the official recorded judgment and two failure‑mode judgments, enabling a comparison between rigid statute matching and ungrounded discretion.
We formalize legal judgment as a reference‑anchored task, meaning a decision must stay tied to both the relevant statute and the factual circumstances. Based on this, we introduce the JusticeAgent framework:
- Element agents extract facts from multimodal evidence;
- Judge agent applies the law guided by skills that encode situational experience.
Skills are distilled from execution trajectories and admitted only under Bayesian credible bounds, guaranteeing that the incorporated experience is statistically reliable.
Experiments reveal a scale‑dependent failure pattern: open‑weight backbones tend to drift toward unsupported reasoning, while frontier models revert to the statutory default. Adding JusticeAgent as a simple plugin to a frozen open‑weight backbone lifts its performance to a commercial level.
All resources, including code and data, are publicly available at https://github.com/beita6969/JusticeAxis.
Review