Banger paper from Google Cloud AI Research.
If you follow autonomous research agents, this one is worth your time.
ScientistTwo takes a problem from a human expert and runs the full discovery cycle without further intervention.
It establishes state-of-the-art baselines, generates seed ideas aimed at resolving stated limitations in existing work, screens each idea on a data subset before committing to full experiments, runs its own ablation studies to attribute the gain, and revises the idea from that attribution.
Manuscript drafting includes a simulated peer-review and rebuttal engine, plus a meta-review pass that feeds back into the work.
They benchmark it against papers accepted at ICLR, ICML and NeurIPS. The reported solutions outperform the human state-of-the-art models on those problems, and the generated papers score higher average ratings than the human-authored ones under automated AI reviewers.