跳到正文
原文
Thomas Wolf· @Thom_Wolf · X·· 2 小时前AI 评分48
AI 导读

人们感到担忧,因为这种作弊率骤降最可能的解释是评测意识:最新的 Opus 模型或许已经聪明到能识别出这个基准测试是在考察作弊行为,并据此调整表现。 若果真如此,这个基准测试就不再能衡量模型"自然"的作弊倾向了。

正文

People are worried because the most likely explanation for such a sudden drop in cheating is evaluation awareness: the latest Opus models may now be smart enough to recognize that this benchmark tests for cheating, and behave accordingly.

If so, the benchmark no longer measures the models' "natural" tendency to cheat.

引用Lukas Petersson@lukaspet
Claude suddenly stopped cheating.
在 X 查看被引用的帖子

来源:Thomas Wolf · x.com