Humans need only a handful of well‑written textbooks to master a field and tackle its hardest problems. We argue that an ideal self‑evolution method should share this property: autonomously learning from raw training material and acquiring transferable problem‑solving ability. Currently there is no direct metric for this capability. StudyBench is introduced as a controlled physics benchmark that measures how efficiently a self‑evolution method converts training material into capability. The test set is split into an Application Set of difficult textbook problems to assess absorption, and a Transfer Set of olympiad‑level problems to assess transfer. Benchmarking representative self‑evolution methods on three base models shows that gains on the Application Set rarely translate to the harder Transfer Set. A guidance ablation reveals a Guidance Gap: even the strongest method captures only a small fraction of the improvement that the same material provides when supplied as in‑context guidance. Moreover, every method hits a Compute Plateau, saturating well before the compute budget is exhausted. Hence the remaining gap is a methodological issue rather than a data or compute problem. By offering a clean, controlled benchmark, StudyBench turns self‑evolution progress from an open‑ended pursuit into a measurable target for future work. The code is released at https://github.com/thunlp/StudyBench.
Review