When one language model judges another model’s code, it often returns a confident verdict with reasoning but does not indicate the lack of supporting evidence. Such a verdict is indistinguishable from one that is truly grounded.
Multi‑agent verification tackles this by decomposing a judgment into checkable claims and seeking independent evidence for each claim. The approach works well when the evidence consists of retrieved documents because the documents naturally satisfy two requirements: (1) they are independent of the answer under review, and (2) they differ between the two candidate solutions being compared.
In code‑judging scenarios, the second requirement frequently breaks down: retrieved code snippets or execution logs tend to be the same for all candidates, providing no discriminative power.
To test this hypothesis, we ran the MARCH framework (unchanged from the original publication) on 80 condition‑by‑cell measurements across two code‑judging benchmarks. MARCH declared both solutions equally good in 78%–95% of pairwise comparisons, achieving only 4.4% accuracy, whereas the same model queried directly on the same tasks reached 43.7% accuracy. Neither easier problems nor a larger judge model improved the result.
We then extracted two label‑free metrics from the pipeline’s own logs. One of these metrics reliably signals when the judge lacks a basis for a decision. By gating on this metric, the pipeline declines comparisons it cannot make, raising accuracy from 20.7% to 36.9% while still answering roughly half of all comparisons.
The main contribution is not a more accurate code judge but a label‑free way to detect when a judge has no grounding, thereby preventing it from guessing.
Review