NeFut Logo NeFut
中 Admin Login

[CS.AI] Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

LLM judges are increasingly employed to decide whether AI outputs satisfy workplace requirements. Agreement on response ordering alone does not guarantee reliable acceptance rates or occupational aggregates. We built the O*NET‑BENCH audit suite from 45,796 worker ratings and tested 33 existing judge configurations across six model families on 4,501 test ratings.

Twenty‑five configurations reach a tie‑aware pair accuracy of at least 0.60, yet a response‑only TF‑IDF baseline fitted on training data nearly matches the strongest judge. Despite ordering agreement, judges estimate that 3.0%‑97.9% of responses are acceptable, compared with 61.1% for occupation‑matched workers.

In one fine‑tuned lineage, switching from pointwise scoring to a few‑shot/listwise protocol improves ordering while reducing agreement with worker means at both task and occupation levels; this reversal replicates on a task‑ and worker‑disjoint validation split under pre‑specified criteria.

Cross‑validated calibration largely removes mean bias, but calibrated scores explain at most 8.5% of individual worker‑rating variance. Prediction‑assisted estimation yields only marginal precision gains at the studied label budgets.

These findings demonstrate that ranking agreement alone is insufficient for occupational measurement. Judges should be validated against acceptance rates and the aggregates their scores will be used to estimate.

Review

Original Source: https://arxiv.org/abs/2610.02492

[h] Back to Home