NeFut Logo NeFut
Admin Login

[CS.AI] Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

Published at: 2026-09-08 22:00 Last updated: 2026-09-09 09:08
#Machine Learning #LLM #Artificial Intelligence

Reference-based automatic evaluation is pivotal for measuring the quality of natural language generation systems. Existing meta-evaluation mainly reports agreement with human judgments or benchmark labels, offering limited insight into how evaluators behave under controlled manipulations.

To address this gap, we introduce a behavioral correctness assumptions framework as a complementary diagnostic for reference-based evaluators. The framework categorizes assumptions into correctness-preserving and correctness-altering groups, and operationalizes each through controlled response transformations that define expected scoring trends.

Controlled transformations encompass lexical substitution, character perturbation, semantic paraphrasing, and other manipulations. Each transformation is paired with a predicted scoring behavior: for instance, a lexical swap that leaves meaning intact should leave scores unchanged under a correctness-preserving assumption, whereas a semantic degradation should trigger a score drop under a correctness-altering assumption.

Our experiments span lexical, character-level, semantic, large language model (LLM)-based, and hybrid evaluators. We assess them on assumption-level behavior, overall stability, sensitivity to input changes, repeat-run variability, configuration sensitivity, and reproducibility.

Results reveal distinct behavioral trade-offs across evaluation paradigms: no evaluator satisfies all proposed correctness assumptions, and evaluators with similar aggregate scores can exhibit markedly different behavioral profiles, differing in assumption compliance and stability.

These findings demonstrate that behavioral correctness assumptions expose diagnostic information hidden from conventional aggregate meta-evaluation, guiding more informed selection of evaluation metrics.

Review: The assumption-based framework adds a fine-grained lens to evaluator analysis, and should be incorporated into future evaluation standards.

Original Source: https://arxiv.org/abs/2609.05289

[h] Back to Home