Evaluating large language models becomes harder as their abilities improve: benchmarks quickly saturate, public test sets risk contamination, and harder tasks often need costly grading or execution infrastructure. These issues are amplified in automatic prompt optimization (APO), where the search for better prompts requires repeated evaluation. To address this we built a chess benchmark from 1,118 Lichess puzzles for studying APO on frozen LLMs, i.e., optimizing prompts without changing model weights.
The chess benchmark offers three key benefits:
- Exact‑match scoring makes evaluation cheap and deterministic;
- Engine‑based analysis of alternative moves provides fine‑grained feedback;
- Problems are renewable with adjustable difficulty, preserving headroom as models improve.
We evaluated six APO algorithms on eight target models, measuring baseline strength, each model's responsiveness to optimization, cross‑model transfer of optimized prompts, and performance in short gameplay rollouts. The strongest model, Gemini 3.5 Flash (used as the meta‑model), solved only about 55% of puzzles, confirming the benchmark’s difficulty. The entire study cost roughly $800, demonstrating affordability.
All puzzles, optimization and evaluation code, and renewal scripts are released on GitHub: https://github.com/imec-ailabs/Automatic-Prompt-Optimization-with-Chess.
Review