Artificial Analysis 发布 Cyber Index 网络安全评估联盟
Announcing the Artificial Analysis Cyber Index Alliance
Artificial Analysis 推出 Cyber Index,整合 CWE-Bench-AA、DeepsecBench-AA 和 CyberGym-E2E-AA 三项评估,覆盖从代码审计、漏洞发现到复现与修补的防御全流程。该指数基于 Stirrup 开源代理框架运行,重点测试模型在拥有源代码访问权限下的漏洞识别与修复能力,并单独记录因安全原因拒绝任务的情况。目前暂不包含利用漏洞生成攻击载荷的能力评估。
Introducing the Artificial Analysis Cyber Index
The Cyber Index Alliance
How the Artificial Analysis Cyber Index works
Benchmark overview
At launch, the Artificial Analysis Cyber Index combines three cybersecurity evaluations from industry partners and academic research. Between them they cover the full breadth of the defensive loop, from scanning a codebase for weaknesses to reproducing a crash and implementing a patch.
| Evaluation | What it measures | Capabilities | Sub-capabilities |
|---|---|---|---|
CWE-Bench-AA Collinear AI | Auditing a real open-source repository for a vulnerability in a described area of concern, then patching it without breaking legitimate behavior. 120 held-out tasks covering all ten OWASP Top 10 (2025) categories. | Identifying and remediating vulnerabilities |
|
DeepsecBench-AA Vercel | Finding vulnerabilities in open-source application code, scored against a golden set of expert-verified findings. | Identifying vulnerabilities |
|
CyberGym-E2E-AA Berkeley RDI | Discovering, reproducing, and patching memory-safety vulnerabilities in C/C++ open-source projects. 131 tasks, one per project, drawn from the 920-instance CyberGym-E2E dataset. | Identifying and remediating vulnerabilities |
|
We tag each evaluation with the capabilities and sub-capabilities it tests, as in the table above, so we can see which parts of cyber defense the Index covers and where it has gaps. It covers identifying and remediating vulnerabilities with access to the source code. We plan to add incident response, writing new code without introducing vulnerabilities, and targets without source access, such as compiled software and live servers. Exploit realization, turning a found vulnerability into a working exploit, is out of scope for a defense-focused index.
Methodology
All three evaluations run on Stirrup, our open-source agent harness.
Because cyber work is dual-use, we also track cases where a model or provider declines a task on safety grounds, and report them separately from the score. Given the open-ended nature of the tasks, the harness in CyberGym-E2E-AA allows models to end a task without a finding if they conclude they cannot find or demonstrate a vulnerability.
Full details are on the methodology page.
CWE-Bench-AA: audit and patch real repositories
DeepsecBench-AA: find the vulnerabilities expert reviewers confirmed
CyberGym-E2E-AA: discover, reproduce, and patch memory-safety bugs
Artificial Analysis Cyber Index resources
Our development roadmap
来源:Artificial Analysis 完整文章(网页) · artificialanalysis.ai