# 公共AI基准测试现状堪忧

- 来源：Ethan Mollick (@emollick)
- 发布时间：2026-09-16 22:40
- AIHOT 分数：48
- AIHOT 链接：https://aihot.news/items/cmu4881ws0d9sro4w1lo7441z
- 原文链接：https://x.com/emollick/status/2100233342494925257

## AI 摘要

公共AI基准测试的现状堪忧，正在削弱我们理解AI当前水平的能力。

大多数知名指标已经饱和，而正如这篇论文所示，尚未饱和的基准测试充斥着大量错误，以至于严重低估了AI的能力。

## 正文

The state of public AI benchmarking is dire and is undermining our ability to understand how good AI is now.

Most famous measures are maxed out, and, as this paper shows, the non-saturated benchmarks are riddled with so many errors that they vastly underestimate AI abilities.
