NeFut Logo NeFut
中 Admin Login

[CS.AI] Agent Evaluation Reliability: More Tasks Won’t Always Fix an Agent Leaderboard

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Agent evaluations are increasingly used to compare large language models and guide deployment, yet leaderboard ranks reflect both the model and the evaluation conditions such as scaffolds or tasks. Consequently, reliability depends on the claim being supported: an evaluation that reliably orders deployed systems may not reliably order the underlying models.

To clarify which conclusions are trustworthy, the authors introduce a Bayesian variance‑decomposition framework tailored for sparse, imbalanced agent leaderboards. They apply it to 22 benchmarks from the Holistic Agent Leaderboard and Harbor Index, separating signal (true model differences relevant to the claim) from noise (irrelevant variation that can still shift rankings).

Key findings are: 1) Reliability hinges on the measurement goal. Fixed model‑scaffold systems achieve high ranking reliability (0.935‑0.994), whereas underlying‑model reliability is markedly lower (0.148‑0.841). 2) Scaffold choice can alter conclusions; inter‑scaffold reliability measures whether scaffolds preserve model rankings, revealing substantial scaffold effects across evaluations. 3) Adding more tasks cannot resolve all uncertainty. When uncertainty is dominated by limited scaffold coverage, even infinitely many similarly constructed tasks improve model‑ranking reliability by at most 0.097. 4) Pooling diverse benchmarks can boost cross‑task ranking reliability from 0.44 to 0.75 under the same task budget and reduce projected cost by up to 83%.

Thus, evaluation design should follow the intended claim: define what a score or ranking means, diagnose factors limiting reliability, and allocate the evaluation budget to the sources of uncertainty that truly matter.

Review

Original Source: https://arxiv.org/abs/2610.00651

[h] Back to Home