# GPT-6 Astra 登顶 TRACES 科学发现基准

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-09-13 01:18
- AIHOT 分数：45
- AIHOT 链接：https://aihot.news/items/cmtyoag8l041trog0zwyoz895
- 原文链接：https://x.com/rohanpaul_ai/status/2098823471480685036

## AI 摘要

Apodex AI 发布 TRACES 基准，从 Tools、Repair、Alternatives、Coherence、Evidence、Scope 六个维度评估 AI 在困难科学问题上的发现能力，而非只给一个总分。

## 正文

A benchmark for difficult scientific discovery from @tianqiao_chen now shows GPT-6 Astra improving across all six capabilities.

But fixing its own mistakes improved the least, Repair improved far less than the other dimensions.

For AI agents, a correct final answer can hide a weak investigation.

An agent might reach the right result after ignoring contradictory feedback. That behavior could fail on the next task, even though the current answer passes its check.

@Apodex_AI released TRACES, a new benchmark for testing AI systems on difficult scientific problems where the correct answer may not already be known.

TRACES separately measures Tools, Repair, Alternatives, Coherence, Evidence, and Scope instead of reducing discovery capability to one score.

Instead of scoring only the final answer, TRACES also evaluates how the AI works through the problem, including its tool use, error correction, evidence, and reasoning process.

### 引用推文

> Apodex：A generational leap, but not on every dimension. Yesterday we put the TRACES board up. Here is the first comparison worth pulling out of it. GPT-6-astra @OpenAI...
