NeFut Logo NeFut
中 Admin Login

[CS.AI] Benchmarking Prompt Optimization of Large Language Models with Chess

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #optimization

Evaluating large language models becomes harder as their abilities improve: benchmarks quickly saturate, public test sets risk contamination, and harder tasks often need costly grading or execution infrastructure. These issues are amplified in automatic prompt optimization (APO), where the search for better prompts requires repeated evaluation. To address this we built a chess benchmark from 1,118 Lichess puzzles for studying APO on frozen LLMs, i.e., optimizing prompts without changing model weights.

The chess benchmark offers three key benefits:

We evaluated six APO algorithms on eight target models, measuring baseline strength, each model's responsiveness to optimization, cross‑model transfer of optimized prompts, and performance in short gameplay rollouts. The strongest model, Gemini 3.5 Flash (used as the meta‑model), solved only about 55% of puzzles, confirming the benchmark’s difficulty. The entire study cost roughly $800, demonstrating affordability.

All puzzles, optimization and evaluation code, and renewal scripts are released on GitHub: https://github.com/imec-ailabs/Automatic-Prompt-Optimization-with-Chess.

Review

Original Source: https://arxiv.org/abs/2610.00416

[h] Back to Home