Benchmark scores are increasingly shaping the development, marketing, and selection of large language models (LLMs). Yet an overall score only makes sense relative to the specific system tested, the questions used, and the evaluation conditions. This article outlines five interrelated limits of general LLM rankings: a gap between evaluated and publicly available systems, commercial incentives and dependencies in external evaluations, benchmark saturation, defective tests and data contamination, models gaming scoring procedures, and the limited relevance of general scores to users' tasks. Documented examples show why each problem calls for a distinct response. We advocate evaluation procedures that disclose the tested configuration, validate the questions and successful task completion, report performance together with cost and execution time, and make the scope of generalization explicit. The paper then discusses Isotanta, a crowdsourced benchmarking platform, as a concrete example of contributed questions and repeated evaluation. A larger question pool can improve task coverage, and repeated sampling can increase the stability of estimates on that pool, though neither guarantees validity or personalization. The platform’s current shared ranking is distinguished from the proposed task‑specific and user‑provided evaluations. The central argument is that model selection requires evidence of performance on the intended work, not merely a high position on a general leaderboard.
Review