Synthetic Sciences just launched its open-source OpenScience agent now outscores Codex and Claude Code on agentic science benchmarks.
A free Apache 2.0 workbench for any provider's model.
On Terminal-Bench-Science, 70 research workflows hosted by Stanford and the Laude Institute, OpenScience solved 53 tasks for 75.7%.
Codex with GPT-6 Astra scored 68.1%, the top entry on a September 23 mirror of the public leaderboard.
The gap widened on Terminal-Bench 4.0's 14 science tasks, where OpenScience scored 71.4% and Claude Code on Claude Fable 5.1 reached 60.0%.
GitHub link in comment