对AI模型的红队评估支持某些主张而不支持另一些主张,而这两者之间的界限是可以计算的,而非仅仅是判断问题。我们将评估的证据上限定义为:在固定测试预算下,一个结果能使信念发生变动的最大倍数;我们针对基准零结果给出了其闭式解,并利用它精确定位了这一界限。
我们发现,在某个可计算的危害率之上,一个规模适中的基准就能以既定的证据标准对某一类别进行认证,此时“零失败”是两种可能观测结果中更强的一种,其权重超过单次可复现的失败。在该比率之下,任何规模可行的被动基准都无法在固定评分规则和近似独立的试验结构下提供所规定的安全性证据。
两种机制之间的交叉点具有闭式解。这一界限并非基准所特有:以程序假设的条件诱发率来表述,它同样涵盖自适应和自动化红队测试,并表明决定证据价值的是假设之间的区分度,而非攻击成功率。我们对照这一界限审计了八套评估体系,发现当前基准对于高频危害类别是充分的,但对于罕见、灾难性危害则相差数个数量级。
安全性基准并非毫无信息量。它们所提供的信息针对的是特定且可计算的一组命题,而它们所需要的纪律是明确说明这些命题是什么。
Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure.
Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones.
Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.