Beautiful paper from Google.
ScientistTwo shows another progress of recursive self-improvement in AI research: it can improve a human method, then use its own discovery as the baseline and improve it again.
the research assistant becomes the research loop: it autonomously improved 86 of 107 human-defined ML problems, suggesting experimentation can be automated long before scientific judgment can.
Instead of helping with one task, ScientistTwo runs a loop: propose ideas, test them, remove what does not help, and use reviewer feedback to start another round.
Across problems based on ICLR, ICML, and NeurIPS papers, it reports an 80.4% success rate and a 25.2% average relative improvement over the original human baselines.