NeFut Logo NeFut
Admin Login

[CS.AI] StudyBench: Can Self‑Evolution Compress Textbooks for Olympiad Capability?

Published at: 2026-09-03 22:00 Last updated: 2026-09-04 02:14
#AI #Machine Learning #LLM

Humans need only a handful of well‑written textbooks to master a field and tackle its hardest problems. We argue that an ideal self‑evolution method should share this property: autonomously learning from raw training material and acquiring transferable problem‑solving ability. Currently there is no direct metric for this capability. StudyBench is introduced as a controlled physics benchmark that measures how efficiently a self‑evolution method converts training material into capability. The test set is split into an Application Set of difficult textbook problems to assess absorption, and a Transfer Set of olympiad‑level problems to assess transfer. Benchmarking representative self‑evolution methods on three base models shows that gains on the Application Set rarely translate to the harder Transfer Set. A guidance ablation reveals a Guidance Gap: even the strongest method captures only a small fraction of the improvement that the same material provides when supplied as in‑context guidance. Moreover, every method hits a Compute Plateau, saturating well before the compute budget is exhausted. Hence the remaining gap is a methodological issue rather than a data or compute problem. By offering a clean, controlled benchmark, StudyBench turns self‑evolution progress from an open‑ended pursuit into a measurable target for future work. The code is released at https://github.com/thunlp/StudyBench.

Review

Original Source: https://arxiv.org/abs/2609.00787

[h] Back to Home