Artificial Analysis 发布 CyberGym-E2E-AA 基准评测,测量模型在内存安全任务上从漏洞发现到打补丁的网络防御能力。部分前沿模型在 85% 以上任务上被安全拦截无法响应;GPT-6 Luna 和 MiMo-V2.6-Pro 在 1M+ 行代码库上跑约 100 次漏洞挖掘仅需约 $20,单任务成本比次强的 Grok 4.7 便宜最多 100 倍。
Restricting offensive without blocking defensive is difficult, as finding and proving a vulnerability requires the same steps whether the goal is to exploit it or patch it
On CyberGym-E2E-AA, a benchmark that measures cyber defense capabilities from discovery to patching on memory-safety tasks, some frontier intelligence models are safety blocked from responding on 85%+ of tasks.
The good news is that some of the most capable models are also the most cost effective. With GPT-6 Luna or MiMo-V2.6-Pro, you can run ~100 bug hunts in a 1M+ line codebase for ~$20 - up to 100x cheaper per task than the next most capable model Grok 4.7.
来源:Artificial Analysis · x.com