What if the ceiling on self-evolution is not the data or the compute, but the loop itself?
Introducing StudyBench: a controlled physics benchmark that measures how efficiently self-evolution converts a fixed corpus of textbooks into capability. It separates what a method absorbs from what it can transfer, turning an open chase into a measurable target — is the bottleneck the data, the compute, or the loop? ✨ Paper: https://arxiv.org/abs/2609.00787 💻 GitHub: https://github.com/thunlp/StudyBench