
Image: METAL
Summary
- On October 8, Anthropic launched the Anthropic Cyber Mission, which combines the Critical Infrastructure Defense Program (CIDP) with OSS Scanner, a free scanning service for open-source projects.
- Anthropic found more than 29,000 candidate vulnerabilities over six months but has had humans review only about 6,000, so it will now send model-generated reports straight to maintainers who opt in, without human review.
- In validation, 85 of 97 critical and high-severity findings (88%) met the disclosure bar with only one false positive, and wolfSSL said 72 of the 74 reports it received were valid and five became CVEs.
Anthropic on October 8 (local time) launched the Anthropic Cyber Mission, a long-term effort that ties together the defense of critical infrastructure and open-source software. Its first two pillars are the Critical Infrastructure Defense Program (CIDP), which pairs frontier models and on-site engineers with the security firms that protect the operational technology (OT) behind power grids, water systems and transportation networks, and OSS Scanner, which gives open-source projects free, periodic scans from the company's strongest models. OSS Scanner's reports go to maintainers exactly as the model wrote them, without human review. Anthropic said it expects a true-positive rate above 90%.
The starting point is a backlog of unprocessed vulnerabilities. According to Anthropic's Frontier Red Team, the company used its latest models over the past six months to scan major software around the world and found more than 29,000 candidate vulnerabilities, but humans have manually reviewed and triaged only about 6,000 of them. The Red Team wrote that "we remain bottlenecked on our human capacity to validate these findings." Maintainers who received first reports increasingly asked for every unverified report along with proposed patches, and nearly 5,000 such reports have been sent so far. OSS Scanner turns that request into a standing service.
Model capability is the basis for that decision. According to the Red Team, the share of vulnerabilities that language models found on CyberGym, an academic vulnerability-discovery benchmark, jumped from under 20% at the beginning of last year to over 85% this year. As a result, the Red Team said, the AI reports open-source maintainers receive have shifted from mostly slop to high-quality bug reports. Reports are generated by the strongest models, including Claude Mythos. The service was inspired by Google's open-source fuzzing service OSS-Fuzz, and unlike Claude Security, the company's paid enterprise product, Anthropic covers the full cost.
Accuracy was measured in advance. Expert penetration testers who review Anthropic's coordinated vulnerability disclosure (CVD) findings checked 97 critical and high-severity vulnerabilities that an early version found across 48 projects, and 85 (88%) met the bar for CVD disclosure. Of the remaining 12, 11 were real bugs that duplicated known issues or other findings from the same scan, and only one was a false positive. Trials with dozens of projects over the past several weeks produced hundreds of bug reports, and several of the vulnerabilities could be chained into unauthenticated remote code execution exploits.
Trial participants gave concrete assessments. Todd Ouska of wolfSSL, an embedded cryptography library, said "of the 74 reports we received, all but two were valid, and five became CVEs," adding, "We'd love more." Anton Arapov of OpenSSL Corporation said "early AI reports about 18 months ago were appalling," while "the reports we received from Anthropic, raw model output included, were as good and sometimes better than what we get from people." Noah Misch of PostgreSQL said several reports came with fixes that could be used nearly as-is, and that fast-track access let the project address the newest issues before a general release.
Each report contains a self-contained reproducer, an explanation of the vulnerability, where possible a bisection tracing when the bug entered the code history, and a candidate patch. According to the OSS Scanner documentation Metal reviewed, scanning agents run only inside hardened sandboxes with internet access fully disabled. Maintainers apply by opening a pull request that adds a project.yaml config file to the GitHub repository anthropics/oss-scanner; the repository URL, a primary contact and a Dockerfile that builds the project offline are required. The eligibility test matches OSS-Fuzz's "critical impact on infrastructure and user security," with decisions made case by case.
The disclosure rules differ from the usual approach. According to the documentation, unvalidated reports carry no deadline such as a 90-day disclosure period. The reasoning is that Anthropic cannot comfortably force maintainers onto a deadline for findings it has not carefully read itself. If a human later validates a report, it may be disclosed under the CVD policy starting 90 days after the maintainer is notified. Maintainers can pause reports by adding the single line disabled: true to the config file. During trials, some maintainers said severity ratings were inflated or that the scanner misunderstood a project's threat model, and Anthropic said it would refine the system based on maintainer feedback.
The critical infrastructure program has 11 founding partners: Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation. The list spans consulting firms, security vendors and equipment manufacturers. According to reports, Andrew Turner, president of commercial cyber at Booz Allen, called operational technology "the next frontier for autonomous AI-enabled attacks." The same reporting noted that the announcement left out commercial terms such as whether partners get free model access and who pays for compute.
Metal previously reported that Anthropic now runs Claude Security scans on Mythos 5, and also covered Anthropic's analysis of GLM-5.3's cyber capabilities. Anthropic said it merged Project Glasswing into its expanded Cyber Verification Program earlier this week. According to reports, Glasswing had given vetted organizations access to Mythos since April. A cyber defense program for state and local governments, launched in June, now provides models and technical support to more than half of US states. The money that keeps OSS Scanner free comes from the Defender Advantage Fund (0xDAF), created in August. Anthropic has also funded the Python Software Foundation, the Apache Software Foundation, and Alpha-Omega and OpenSSF under the Linux Foundation.
Anthropic's launch post on its official X account, which Metal checked, drew about 138,000 views and some 1,500 likes in less than a day. From a lawyer's perspective, the core of this design is that it moves where responsibility sits. In exchange for dropping human review, it sets no disclosure deadline and lets maintainers decide enrollment and opt-out with a single pull request. In effect, it accepts only projects that can bear the cost of false positives. Anthropic forecasts that AI will favor defense within two years, but said that for now the cost of exploiting vulnerabilities has fallen while verifying and fixing them is still slow and depends on people. In operational technology, a patch must wait until it can be applied safely to running equipment, which in rare cases can take decades.
Sources
- Anthropic — Introducing the Anthropic Cyber Mission →
- Anthropic Frontier Red Team — Launching an opt-in vulnerability-finding service for open-source software →
- Anthropic — OSS Scanner →
- The Verge — Anthropic launches free AI security scans for open-source projects →
- SiliconANGLE — Anthropic launches critical infrastructure program and free OSS Scanner for open source →





Comments