NeFut Logo NeFut
Admin Login

[CS.AI] What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

Published at: 2026-09-18 22:00 Last updated: 2026-09-20 12:54
#Machine Learning #LLM #Artificial Intelligence

This paper systematically surveys 14,767 arXiv papers that introduced or updated evaluation resources between January 2022 and August 2026, revealing how benchmark design for large language models (LLMs) has evolved. Using staged screening and automated full‑text coding, we examine shifts in target systems and domains, evaluation materials and conditions, and scoring mechanisms. Findings show a move from pure text understanding toward action, interaction, and professional applications; classic metrics such as accuracy coexist with newer ones like operability and user satisfaction. Model participation is uneven: LLM‑based scoring grows in both agent and non‑agent benchmarks, while model‑generated test items do not show a sustained rise. The study raises the question of whether AI‑driven test construction, task execution, and judgment provide more independent evidence or merely reproduce the preferences and blind spots of the models themselves.

Review

Original Source: https://arxiv.org/abs/2609.19182

[h] Back to Home