NeFut Logo NeFut
中 Admin Login

[CS.AI] Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

Pairwise language‑model judges can gather evidence by direct comparison, reasoning, or reference‑based verification, yet no single protocol dominates across benchmarks and judge backbones. We introduce Backbone‑Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference direction but cannot alter its magnitude. BAER separates each expert's signed preference from candidate‑invariant reliability and builds three symmetric heads—evidence stacking, reliability‑based expert routing, and candidate‑blind reference verification. Development data select one head for each benchmark‑backbone condition, and this choice is frozen before testing. Across four benchmarks and two 8B judge backbones, BAER attains the highest test accuracy in all eight conditions, with full prediction coverage and gains of 0.87‑7.32 points over the strongest external baseline. The results demonstrate that adapting how evidence is gathered is more reliable than fixing a single judging protocol everywhere.

Review: BAER’s dynamic evidence routing yields consistent accuracy improvements across diverse benchmarks and model backbones, confirming the practical value of adaptive evidence selection.

Original Source: https://arxiv.org/abs/2609.30751

[h] Back to Home