LLM-as-a-judge has become the de‑facto standard for large‑scale, subjective evaluation. Current leaderboards try to offset systematic measurement bias by running ever more pairwise comparisons, which is statistically unsound and computationally wasteful. The root problem is an incomplete measurement model that treats LLM judges as neutral, interchangeable instruments, ignoring documented biases such as position bias, verbosity bias, judge severity, and self‑enhancement. These biases cannot be eliminated simply by collecting more data.
We introduce a unified latent‑variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders. The approach yields reliable rankings with far fewer comparisons. Because fitting the model costs negligible compute relative to a single round of LLM inference, bias correction is not only statistically rigorous but also a more sustainable path to trustworthy evaluation.
Review