Language‑model judges now filter training data, score generations, and drive leaderboards, effectively becoming measurement instruments. The instrument rests on a rarely stated assumption: the same request sent to the same model name will produce the same output tomorrow. We audited this assumption in two preregistered campaigns, fixing every threshold in advance; neither campaign validated its instrument.
Across 52,988 audited request attempts, same‑window repeat rankings achieved a Spearman correlation of only 0.400 against a required 0.90, and byte‑identical next‑day replays yielded 0.78 against a required 0.99, each time with execution records at ceiling. Three mechanisms explain the gap: (1) a label‑to‑meaning mapping that biases readouts as strongly as the signal; (2) candidate gaps seven orders of magnitude below the instrument’s own noise floor; (3) byte‑identical inputs returning different rankings, a noise that exact‑permutation readouts compound. Neither metric substitution nor resampling repaired the issue on the tested grid.
Preregistered follow‑ups bounded the problem: waiting did not help on sampled days (0.805 vs. 0.800, replicated over five further days); switching providers did not help (four providers shared the floor, medians 0.74‑0.88, none explained by exposed metadata); self‑hosting on batch‑invariant kernels helped only while the server was quiet; and constructed errors with known gaps showed that the readout’s separation tracks error type, not size.
We distilled the evidence into a three‑level snapshot‑identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study’s call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.
Review