Direct-decision models map text to low‑latency structured labels and scores, making them attractive for classification and automatic evaluation. Reliability, however, demands more than accuracy; a model must faithfully employ the ordinal decision scale supplied by the user.
We evaluate JEV 1.13 and three open KEV models. Starting with ANLI, JEV assigns 38.8% of all predictions and 51.3% of errors to Neutral despite a 74.95% accuracy, nearly balanced gold labels, and balanced candidate positions.
Across 36 ordinal datasets, final decisions use only 67%–76% of the effective gold support, whereas on four nominal tasks utilization reaches 87%–102%. Randomizing candidate order weakens but does not eliminate this compression.
Holding items and source scores fixed while balancing gold support and positions, we vary the scale size $K$ from 2 to 14; utilization drops for every model, falling to 26%–75% at $K=14$, even though most models retain broad candidate probability distributions.
Targeted BA‑LoRA post‑training raises gold‑relative utilization from roughly 47% to 86% on eight supervised scales for both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit.
We call this phenomenon "ordinal scale‑utilization bias": a decision‑stage candidate‑space compression distinct from accuracy, gold imbalance, fixed position, or candidate count alone.
The code and data are released at https://github.com/Glax147/jev_ordinal_scale_bia.
Review