Thariq · @trq212 · X·2026-09-12 03:13·52分钟前
AI 导读

如今仅凭通过/失败分数基本无法解读评测结果 我在基准测试中看到的许多失败,都源于过于严格的隐藏测试,某些情况下模型的答案比预期的评测结果更合理

Thariq@trq212
41AI 编辑部评分,满分 100
2026-09-12 03:13· 52分钟前
AI 导读

如今仅凭通过/失败分数基本无法解读评测结果 我在基准测试中看到的许多失败,都源于过于严格的隐藏测试,某些情况下模型的答案比预期的评测结果更合理

it's basically impossible to interpret evals by looking at just at the pass/fail scores these days

many of the failures I see in benchmarks are due to overly strict hidden tests, in some cases the model's answer makes more sense than the expected eval result

来源:Thariq· x.com