Reference‑based LLM‑as‑a‑judge evaluation assumes the reference answer is the target. In deployed agentic systems that act on dynamic entities such as support cases, assets, or accounts, the nearest reference usually describes the correct procedure applied to a different entity. Consequently, a literal judge penalizes mismatched identifiers, dates, and statuses as errors or hallucinations, a failure mode we call reference‑instance divergence (RID).
The CARGO framework addresses RID with three mechanisms: (i) treat retrieved references as procedural exemplars and ground factual judgments in the live instance’s observed context; (ii) assign each claim a three‑way status—supported, contradicted, unverifiable—and penalize only contradictions; (iii) gate evaluation by retrieval confidence, casting production evaluation as selective prediction.
To assess CARGO, the authors built CARGO‑Bench, a perturbation‑based diagnostic suite that constructs ground truth by design and separates leniency from discrimination. On 246 items, two judge models, and 7,872 judgments, the standard reference‑based judge penalized 100% of correct entity‑transplanted answers and yielded a discrimination index (DI) near 0; providing live facts without re‑framing changed nothing. CARGO eliminated these false penalties (0/50) while retaining near‑complete contradiction recall (50/50 and 49/50), raising DI to 0.58 [0.48, 0.68]. A rubric‑swap control showed that most of the gain stems from context‑grounded dimension definitions.
CARGO also reveals a limitation: the leniency that protects entity values suppresses detection of procedural corruptions, resulting in only 20% recall for such errors. A post‑hoc fix did not close the gap, and an LLM‑as‑annotator study with written guidelines and adjudication exhibited the same blind spot.
The authors release a preregistered protocol for extending the evaluation to expert agreement, risk coverage, and cost on production traffic.
Review