Improving reasoning abilities of large language models (LLMs) requires high‑quality data that expose difficult decisions, competing alternatives, and their consequences. Data scarcity stems from low‑quality synthetic data and the high cost of human labeling. We introduce Self‑Play Search Distillation (SPSD), a framework that generates superhuman synthetic data by self‑playing MuZero‑like networks on board games.
SPSD leverages executable environments to turn the search process into structured reasoning problems. At each state the expert model identifies a preferred action, plausible alternatives, possible opponent replies, and value estimates. These self‑play search records are converted into superhuman chains‑of‑thought, providing environment‑grounded supervision for training LLMs.
Although trained only on self‑play records, SPSD transfers to unseen mathematics. On Qwen3‑4B‑Base, the mean score across six math benchmarks rises from 24.1 to 36.6 and the held‑out game win rate climbs from 15% to 45%.
SPSD offers an annotation‑efficient way to create high‑quality synthetic data, substantially boosting LLM performance on reasoning tasks.
Review