Artificial Analysis 发布 Terminal-Bench-Science 0.1 榜单

Artificial Analysis · @ArtificialAnlys · X·2026-09-25 07:30·1小时前
AI 导读

Artificial Analysis 上线 Terminal-Bench-Science 0.1 智能体科研基准榜单,GPT-6 Astra (max) 以 63% 居首,Claude Opus 5.5 (xhigh) 以 62% 紧随其后。

Artificial Analysis@ArtificialAnlys
44AI 编辑部评分,满分 100

Artificial Analysis 发布 Terminal-Bench-Science 0.1 榜单

2026-09-25 07:30· 1小时前
AI 导读

Artificial Analysis 上线 Terminal-Bench-Science 0.1 智能体科研基准榜单,GPT-6 Astra (max) 以 63% 居首,Claude Opus 5.5 (xhigh) 以 62% 紧随其后。

Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62%

Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8).

As with Terminal-Bench, each task drops an agent into a sandbox environment with the data, tools and instructions for a task, and the agent runs from start to finish. Automated tests grade every task pass/fail, and we report the average pass@1 over 3 attempts.

Key takeaways:

➤ Only two models score above 50%: GPT-6 Astra (max) at 63%, and Claude Opus 5.5 at 62% (xhigh) and 59% (max). Recent releases have made large advancements, but even the top model, GPT-6 Astra, has significant headroom

➤ Life sciences is the lowest-scoring domain for most leading models, though domain scores are noisy with 8 to 19 tasks each. Claude Opus 5.5 (xhigh) passes 71% of mathematical sciences tasks but 46% of life sciences tasks

➤ The best open weights models, GLM-5.3 (max) at 10% and DeepSeek V4.1 Flash (max) at 9%, sit more than 50 points below the leaders

来源:Artificial Analysis· x.com