NeFut Logo NeFut
中 Admin Login

[CS.AI] Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Benchmarks scored by an LLM judge can discriminate differences as small as a tenth of a point, yet the resolution of such scores has never been quantified. Prior sample‑complexity work only addresses accuracy benchmarks and leaves the judged case open. We treat the 373,019 judgments as measurements of a system and apply generalizability theory to decompose variance into system, item, judge, and system‑judge interaction components. The key structural result is that with a single judge the generalizability coefficient converges to $$\frac{\sigma^2_s}{\sigma^2_s+\sigma^2_{sj}}$$ regardless of the number of items, because the system‑judge term contains no $n_i@@@MATH_BLOCK2@@@\sigma^2{sj}$ drops by two orders of magnitude and the ceiling rises to 0.986 (bootstrap [0.934, 1.000] on 11 systems), making a single judge sufficient. Pairwise evaluation introduces a new problem: a system shown first wins 8.6 percentage points more often than when shown second, a bias 1.23× the median improvement reported in the 53 published win‑rate comparisons we recovered. Protocol design dominates panel size. Measured floors on a 0‑5 scale range from 0.41 to 1.24 points, whereas the median reported improvement is only 0.28 points; on the one benchmark that allowed an exact match, all 17 recovered MT‑Bench improvements fall below MT‑Bench’s own floor, and 70 % of win‑rate claims fall below the pairwise floor. An audit of 628 arXiv papers (double‑coded by two independent models and validated against blind human coding, kappa = 0.73) finds fewer than one in four papers state whether their evaluation was run more than once, and only 46‑67 % report any uncertainty.

Review

Original Source: https://arxiv.org/abs/2609.27787

[h] Back to Home