跳到正文
原文
Artificial Analysis 完整文章(网页)·· 4 天前AI 评分40

Artificial Analysis 发布 Cyber Index 网络安全评估联盟

Announcing the Artificial Analysis Cyber Index Alliance

AI 导读

Artificial Analysis 推出 Cyber Index,整合 CWE-Bench-AA、DeepsecBench-AA 和 CyberGym-E2E-AA 三项评估,覆盖从代码审计、漏洞发现到复现与修补的防御全流程。该指数基于 Stirrup 开源代理框架运行,重点测试模型在拥有源代码访问权限下的漏洞识别与修复能力,并单独记录因安全原因拒绝任务的情况。目前暂不包含利用漏洞生成攻击载荷的能力评估。

正文

Introducing the Artificial Analysis Cyber Index

The Cyber Index Alliance

How the Artificial Analysis Cyber Index works

Benchmark overview

At launch, the Artificial Analysis Cyber Index combines three cybersecurity evaluations from industry partners and academic research. Between them they cover the full breadth of the defensive loop, from scanning a codebase for weaknesses to reproducing a crash and implementing a patch.

EvaluationWhat it measuresCapabilitiesSub-capabilities

CWE-Bench-AA

Collinear AI
Auditing a real open-source repository for a vulnerability in a described area of concern, then patching it without breaking legitimate behavior. 120 held-out tasks covering all ten OWASP Top 10 (2025) categories.Identifying and remediating vulnerabilities
  • Discovering and scanning for vulnerabilities
  • Implementing a patch or mitigation

DeepsecBench-AA

Vercel
Finding vulnerabilities in open-source application code, scored against a golden set of expert-verified findings.Identifying vulnerabilities
  • Discovering and scanning for vulnerabilities

CyberGym-E2E-AA

Berkeley RDI
Discovering, reproducing, and patching memory-safety vulnerabilities in C/C++ open-source projects. 131 tasks, one per project, drawn from the 920-instance CyberGym-E2E dataset.Identifying and remediating vulnerabilities
  • Discovering and scanning for vulnerabilities
  • Reproducing and validating vulnerabilities
  • Implementing a patch or mitigation

We tag each evaluation with the capabilities and sub-capabilities it tests, as in the table above, so we can see which parts of cyber defense the Index covers and where it has gaps. It covers identifying and remediating vulnerabilities with access to the source code. We plan to add incident response, writing new code without introducing vulnerabilities, and targets without source access, such as compiled software and live servers. Exploit realization, turning a found vulnerability into a working exploit, is out of scope for a defense-focused index.

Methodology

All three evaluations run on Stirrup, our open-source agent harness.

Because cyber work is dual-use, we also track cases where a model or provider declines a task on safety grounds, and report them separately from the score. Given the open-ended nature of the tasks, the harness in CyberGym-E2E-AA allows models to end a task without a finding if they conclude they cannot find or demonstrate a vulnerability.

Full details are on the methodology page.

CWE-Bench-AA: audit and patch real repositories

DeepsecBench-AA: find the vulnerabilities expert reviewers confirmed

CyberGym-E2E-AA: discover, reproduce, and patch memory-safety bugs

Artificial Analysis Cyber Index resources

Our development roadmap

来源:Artificial Analysis 完整文章(网页) · artificialanalysis.ai