# 基准测试隐藏测试过严问题

- 来源：Thariq (@trq212)
- 发布时间：2026-09-12 03:13
- AIHOT 分数：41
- AIHOT 链接：https://aihot.news/items/cmtxcys3v02vuro6p8gql6b9c
- 原文链接：https://x.com/trq212/status/2098490139798655427

## AI 摘要

如今仅凭通过/失败分数基本无法解读评测结果

我在基准测试中看到的许多失败，都源于过于严格的隐藏测试，某些情况下模型的答案比预期的评测结果更合理

## 正文

it's basically impossible to interpret evals by looking at just at the pass/fail scores these days

many of the failures I see in benchmarks are due to overly strict hidden tests, in some cases the model's answer makes more sense than the expected eval result
