NeFut Logo NeFut
中 Admin Login

[CS.AI] Game Arena: Strategic Evaluation of LLMs in Competitive Settings

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

We introduce Kaggle Game Arena, an open and continuously expanding platform for evaluating large language models (LLMs) via competitive games. Unlike static benchmarks, the arena lets models face each other in structured settings, where gameplay difficulty naturally rises as models improve, preventing performance saturation.\ \ The technical report outlines the underlying infrastructure and describes three pilot game environments: Chess, Poker, and Werewolf. These cover perfect‑information, imperfect‑information, and multiplayer scenarios, enabling systematic study of models' strategic planning, adaptability, and robustness under uncertainty.\ \ For each game we detail the environment, evaluation metrics, and results from full competitions across multiple models. With robust infrastructure and large‑scale ground‑truth evaluation, Game Arena ensures reproducibility, transparency, and long‑term generalizability to new games and variants.\ \ Review: By embedding LLMs in dynamic competitive contexts, the arena offers a more realistic assessment of strategic reasoning and collaborative abilities, paving the way for continual model advancement.

Original Source: https://arxiv.org/abs/2609.31473

[h] Back to Home