LLM-as-a-Judge pipelines are increasingly employed to assess AI‑generated text, under the assumption that the model reasons over a rubric to judge candidate responses. We scrutinize this assumption and discover that the rubric itself encodes recoverable evaluative signals.
In our experiments, classifiers trained solely on rubric text—without any access to the evaluated responses—achieve significant predictive performance on judge outputs, indicating that scores can be partially anticipated independent of model outputs.
Counterfactual perturbation studies further reveal that when either the candidate response or the rubric criterion is flipped, LLM judges often fail to reliably update their decisions, leading to unchanged or erroneous judgments.
These findings raise concerns about the reliability of rubric‑based LLM evaluation and highlight the need for more systematic methodological investigations of LLM‑driven automated assessment.
Review