# Epoch AI 推出 Benchmark Reviews 计划，首批审计 15 个 AI 基准

- 来源：Karina (@karinanguyen)
- 发布时间：2026-09-18 08:42
- AIHOT 分数：50
- AIHOT 链接：https://aihot.news/items/cmu69kf9r0g2orofjhiqe1qnr
- 原文链接：https://x.com/karinanguyen/status/2100747283557998629

## AI 摘要

Epoch AI 推出 Benchmark Reviews 计划，用于审计 AI 基准，首批覆盖 15 个基准：4 个 Verified、9 个 Flawed、2 个信息不足以评审。Verified 包括 WeirdML v2、ExploitBench v0.1、PostTrainBench v1.1、SimpleQA Verified，Flawed 包括 Terminal-Bench 4.0.0、SWE-Bench Verified、SWE-Bench Pro 等。作者 Karina Nguyen 转发并称赞该举措能激励行业构建真正高质量的基准，并感谢其审计 PostTrainBench 和挖掘旧的 SimpleQA。

## 正文

Great initiative to eval the evals. It incentivizes the industry to build genuinely high-quality benchmarks.

Ty for auditing PTB and digging up the old SimpleQA too :)

### 引用推文

> Epoch AI：Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough infor...
