Ethan Mollick · @emollick · X·2026-09-16 22:40·26分钟前
AI 导读

公共AI基准测试的现状堪忧,正在削弱我们理解AI当前水平的能力。 大多数知名指标已经饱和,而正如这篇论文所示,尚未饱和的基准测试充斥着大量错误,以至于严重低估了AI的能力。

Ethan Mollick@emollick
48AI 编辑部评分,满分 100
2026-09-16 22:40· 26分钟前
AI 导读

公共AI基准测试的现状堪忧,正在削弱我们理解AI当前水平的能力。 大多数知名指标已经饱和,而正如这篇论文所示,尚未饱和的基准测试充斥着大量错误,以至于严重低估了AI的能力。

The state of public AI benchmarking is dire and is undermining our ability to understand how good AI is now.

Most famous measures are maxed out, and, as this paper shows, the non-saturated benchmarks are riddled with so many errors that they vastly underestimate AI abilities.

来源:Ethan Mollick· x.com