Anthropic 推出开源漏洞扫描服务 OSS Scanner,免费面向开源项目
Launching an opt-in vulnerability-finding service for open-source software
Anthropic 发布 OSS Scanner,一个可选加入的开源漏洞扫描服务,用其最强模型(包括 Claude Mythos)定期免费扫描开源项目,输出为全模型生成、无人工复核的报告。
原文给出扫描准确率、人工验证数据和维护者反馈,读者可评估这个免费开源安全扫描服务的实际信号质量与参与方式。
We’re launching OSS Scanner, an opt-in vulnerability scanner for the open-source ecosystem informed by our experience using Claude to find vulnerabilities during Project Glasswing. Projects that join will receive thorough, periodic security scans by our strongest models at no cost.
Language models have rapidly advanced in their ability to discover vulnerabilities, as we have reported extensively on this blog. On CyberGym, one academic vulnerability-finding benchmark, LLMs have gone from finding under 20% of vulnerabilities at the beginning of last year to finding over 85% this year. As a result, open source maintainers have gone from receiving mostly slop from LLMs to receiving high-quality bug reports.
Over the last six months, we’ve used our latest models to scan for vulnerabilities in some of the world’s most important software projects. We have discovered over 29,000 candidate vulnerabilities, but have only been able to manually review and triage approximately 6,000 of these. While 6,000 vulnerabilities is significant, we remain bottlenecked on our human capacity to validate these findings.
We are working on scaling our vulnerability disclosures. Meanwhile, with increasing frequency, maintainers who receive our first reports simply ask us for a bulk submission of all the unverified reports with proposed patches: to date, we have sent nearly 5,000 reports directly to maintainers after they asked to receive everything we had—even if it wasn’t validated. Since exploits can now be developed in minutes, projects that can find and address vulnerabilities faster can better secure their software against attackers who are racing to find and leverage these same weaknesses.
We will continue to manually disclose human-verified vulnerability reports via our existing coordinated vulnerability disclosure (CVD) process, especially for projects without the resourcing to triage reports themselves. But we are now making available an optional fast-track for those who wish to receive vulnerability reports as soon as they become available.
Our vulnerability scanner
Inspired by the positive impact on the open source ecosystem of Google's OSS-Fuzz, a project that scans open-source software for vulnerabilities with fuzzers, we are launching OSS Scanner to find vulnerabilities in open-source code with our strongest language models. While Claude Security, our general-access code scanning and patching product focuses on helping enterprises defend their systems, OSS Scanner provides security audits at no cost to open-source projects.
The outputs of this opt-in vulnerability scanner will be fully model-generated, without human review or triage. This will enable faster and more frequent scanning, but means that it is possible reports will be incorrect or invalid. These reports will be generated by our strongest models (including Claude Mythos) to give open-source projects the largest defensive advantage.
We have spent the last several weeks validating this pipeline with dozens of open source projects. These initial disclosures contained hundreds of bug reports, including multiple vulnerabilities that we were able to chain to unauthenticated remote code execution exploits affecting these projects. Each report contains a self-contained reproducer, explanation of the vulnerability (including a bisection to determine when the bug was introduced, where possible), and candidate patch (when available) for how the bug can be fixed. We are now making this service available to more open source projects.
Some of the feedback we received as we tested early versions of OSS Scanner and our raw model outputs:
- “An unusually high fraction of OSS Scanner’s findings uncovered PostgreSQL defects. Several reports came with fixes we can use nearly as-is, and fast-track access let us address the newest issues before they reached a GA release.” —Noah Misch, PostgreSQL
- “Early AI reports about 18 months ago, before Project Glasswing, were appalling. The reports we received from Anthropic, raw model output included, were as good and sometimes better than what we get from people. Particularly when a report comes with a real exploit attached, that's basically job done for an engineer as you can verify it right away” —Anton Arapov, OpenSSL Corporation
- “We found the signal from these reports high: of the 74 reports we received, all but two were valid, and five became CVEs. With patches attached, the reports slotted right into our existing process to verify and fix issues. We’d love more.” —Todd Ouska, wolfSSL
- “The bug reports were thorough and clear, with a strong understanding of HotCRP's complex permission model and good bug prioritization.” —Eddie Kohler, HotCRP
To validate an early version of OSS Scanner, we asked the expert penetration testers who review our CVD findings to check 97 critical and high-severity vulnerabilities from the scanner across 48 projects. Of these, 85 (88%) met the bar for our CVD process. Of the remaining 12, 11 were real but duplicated known issues or other findings from the scan, and only one was invalid, i.e. a “false positive.” We’ve since sent many of these findings to maintainers, who have seldom told us a high or critical finding was invalid, consistent with our prior open-source work. Some have told us severity ratings can be inflated or the scanner misunderstood the project’s threat model. We can’t guarantee the scanner will be perfect, but we’ll keep refining the system based on maintainer feedback and as models improve.
Getting started
Core maintainers of eligible projects can enroll by submitting a PR to this GitHub repo following the standard project template, with additional guidance provided in our extended FAQ. Projects are eligible based on a similar set of criteria that OSS-Fuzz uses: briefly, projects should have a “critical impact on infrastructure and user security” and we will make decisions on a case-by-case basis.
Whether or not OSS Scanner is right for you, our new Cyber Verification Program makes advanced cyber capabilities and reduced blocking classifiers available to qualifying security professionals. Claude for OSS provides free Claude Max 20x subscriptions to help remediate vulnerabilities and improve OSS projects.
Related content
Claude-shaped science
Guest author Prof. Matthew Schwartz describes what happened when he stopped fighting Claude and allowed Claude to find “Claude-shaped” problems: ones best suited to the capabilities of the current generation of LLM tools. This led him to build BootLoops, a toolkit for exact calculations in quantitative science, which he has been applying across scientific fields alongside experts.
What work can robots do?
We built an index of how well today’s robots can perform US job tasks. Robots can already do three-quarters of physical tasks, mostly in limited settings, but are cost-competitive for just 0.3% of them.
What do you want from AI?
We’re launching a new study using Anthropic Interviewer to learn from your experiences with AI.
来源:Anthropic:Research · anthropic.com