Tidavo
Wednesday 7 October 2026

Anthropic Creates Three Verification Levels to Give Cyber Defenders Access to Powerful Models

The artificial intelligence developer is loosening automated guardrails for vetted specialists who test software and defend vital networks.

7 Oct 2026

Artificial intelligence systems capable of identifying software flaws present an acute double-edged risk. The exact technical reasoning that allows a defensive analyst to patch a zero-day exploit can just as easily help an intruder craft an attack. Because of this overlap, standard public configurations of leading systems automatically reject prompts involving exploit development, malware examination, and penetration testing.

Anthropic is restructuring how it grants exemptions to that automated censorship. The company announced an updated Cyber Verification Program that unifies its previous defensive pilots, including an initiative called Project Glasswing, into a single framework. Under the revised arrangement, vetted practitioners receive tiered clearances to deploy models such as Claude Opus 5.5, Sonnet 5.5, and Mythos 5.1 on security workloads.

The baseline tier, labeled Defense Access, targets standard defensive operations. It permits incident response squads, threat intelligence analysts, and software maintainers to analyze suspicious binaries and confirm system vulnerabilities. Eligibility extends to corporate teams, public utilities, open-source developers, and individual security researchers who hold an established record of responsible disclosure. Reviews at this tier are designed to finish within several days.

A second level, known as Red Team Access, expands capabilities to active penetration testing and simulated intrusions. Anthropic limits this tier strictly to organizations hired or assigned to test authorized targets, explicitly excluding independent researchers. Requests take several weeks to process, though applicants receive defensive permissions while waiting. Hard guardrails stay active in real time to prohibit destructive behaviors, including the deployment of extortion software or actions threatening physical infrastructure.

The most permissive level, Specialized Access, is restricted to entities responsible for safety-critical environments. These include air traffic communications, electrical grids, telecommunications backbones, and interbank transaction networks. Anthropic vets these participants in coordination with the United States government, automatically transferring existing Glasswing members directly into this tier.

Anthropic measured the practical effect of these boundaries using CyScenarioBench, an internal test designed to evaluate multistep cyber operations. On standard accounts, the model hit automated refusals on all 50 test challenges at the first prompt. Under Defense Access settings, it blocked 46 attempts and completed four. In the Red Team tier, the model completed 34 of the 50 challenges without triggering safety refusals.

The push to relax controls follows substantial discoveries during initial testing phases. Organizations collaborating under Project Glasswing identified more than 129,000 verified software bugs between April and July, while the company's internal scans uncovered 5,500 additional flaws in open-source projects through October. More than 33,000 of these defects were classified as critical or severe, with researchers noting that automated assistance reduced review cycles from months to days.

To counter abuse, the program requires participants to consent to prompt monitoring and data logging. Anthropic intends to introduce client-managed cloud storage safeguards later in the autumn, allowing select organizations to isolate sensitive findings within their own digital boundaries. The tiered program is accessible across the Claude Platform, Google Cloud Vertex AI, and Microsoft Foundry, with restricted availability on Amazon Bedrock.

SecurityWeek , Security Affairs , Help Net Security