耶鲁等机构论文:物理基准评测质量差,GPT-5.6-Sol 的 HLE-Physics 修正后从 47.3% 升至 78.7%

Rohan Paul · @rohanpaul_ai · X·2026-09-17 04:25·1小时前
AI 导读

耶鲁大学联合多所机构论文审计 6 个主流物理基准后发现,模型被误判的失败大多源于糟糕题目、错误参考答案和脆弱的评分器。在 4 个被审计基准子集的 250 个被拒案例中,仅 12 个是模型真实错误;GPT-5.6-Sol 经专家复评后 HLE-Physics 分数从 47.3% 升至 78.7%。

Rohan Paul@rohanpaul_ai
58AI 编辑部评分,满分 100

耶鲁等机构论文:物理基准评测质量差,GPT-5.6-Sol 的 HLE-Physics 修正后从 47.3% 升至 78.7%

2026-09-17 04:25· 1小时前
AI 导读

耶鲁大学联合多所机构论文审计 6 个主流物理基准后发现,模型被误判的失败大多源于糟糕题目、错误参考答案和脆弱的评分器。在 4 个被审计基准子集的 250 个被拒案例中,仅 12 个是模型真实错误;GPT-5.6-Sol 经专家复评后 HLE-Physics 分数从 47.3% 升至 78.7%。

New Yale Univ + other top lab paper shows frontier models are already close to maxing out today’s closed-ended physics benchmarks.

And many apparent failures come from bad questions, wrong reference answers, and brittle graders explain most audited physics failures, making benchmark quality the new bottleneck for measuring frontier models.

The researchers had physicists re-check model failures across 6 popular physics benchmarks instead of trusting the original scores.

In 250 rejected cases from 4 audited benchmark subsets, only 12 were actual model mistakes.

The other 238 came from bad questions, wrong reference answers, or graders rejecting correct answers.

After expert review, GPT-5.6-Sol’s measured HLE-Physics score rose from 47.3% to 78.7%.

So a low physics benchmark score can badly underestimate what a frontier model can actually solve.

But that does not mean these models can reliably do physics research: the authors’ agents still failed to fully solve any of the open theoretical-physics problems they tried.

but it does mean that do not treat benchmark scores as clean ground truth anymore for AI's Physics capability.

来源:Rohan Paul· x.com