NeFut Logo NeFut
Admin Login

[CS.AI] Judging LLM-as-a-Judge: Concerns over Rubric Artifacts in Automated Text Generation Evaluation

Published at: 2026-09-05 22:00 Last updated: 2026-09-06 01:02
#AI #Machine Learning #LLM

LLM-as-a-Judge pipelines are increasingly employed to assess AI‑generated text, under the assumption that the model reasons over a rubric to judge candidate responses. We scrutinize this assumption and discover that the rubric itself encodes recoverable evaluative signals.

In our experiments, classifiers trained solely on rubric text—without any access to the evaluated responses—achieve significant predictive performance on judge outputs, indicating that scores can be partially anticipated independent of model outputs.

Counterfactual perturbation studies further reveal that when either the candidate response or the rubric criterion is flipped, LLM judges often fail to reliably update their decisions, leading to unchanged or erroneous judgments.

These findings raise concerns about the reliability of rubric‑based LLM evaluation and highlight the need for more systematic methodological investigations of LLM‑driven automated assessment.

Review

Original Source: https://arxiv.org/abs/2609.02942

[h] Back to Home