# Artificial Analysis 发布 Terminal-Bench-Science 0.1 榜单

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-09-25 07:30
- AIHOT 分数：44
- AIHOT 链接：https://aihot.news/items/cmug6quo10h9frogvgzlyf4o8
- 原文链接：https://x.com/ArtificialAnlys/status/2103265956479070260

## AI 摘要

Artificial Analysis 上线 Terminal-Bench-Science 0.1 智能体科研基准榜单，GPT-6 Astra (max) 以 63% 居首，Claude Opus 5.5 (xhigh) 以 62% 紧随其后。

## 正文

Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62%

Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8).

As with Terminal-Bench, each task drops an agent into a sandbox environment with the data, tools and instructions for a task, and the agent runs from start to finish. Automated tests grade every task pass/fail, and we report the average pass@1 over 3 attempts.

Key takeaways:

➤ Only two models score above 50%: GPT-6 Astra (max) at 63%, and Claude Opus 5.5 at 62% (xhigh) and 59% (max). Recent releases have made large advancements, but even the top model, GPT-6 Astra, has significant headroom

➤ Life sciences is the lowest-scoring domain for most leading models, though domain scores are noisy with 8 to 19 tasks each. Claude Opus 5.5 (xhigh) passes 71% of mathematical sciences tasks but 46% of life sciences tasks

➤ The best open weights models, GLM-5.3 (max) at 10% and DeepSeek V4.1 Flash (max) at 9%, sit more than 50 points below the leaders
