One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

OpenAI's Astra flagged for potential "Critical" cyber capability

OpenAI said an internal evaluation released on August 7, 2026 found it could not rule out that its upcoming model, Astra, meets the "Critical" cyber capability threshold under its Preparedness Framework.

이미지: AI 생성 — METAL LAB

Summary

  • OpenAI announced that an internal evaluation of its upcoming model Astra could not rule out that it meets the "Critical" cyber capability threshold under the Preparedness Framework.
  • Previous models, including GPT-5.6-Sol, were rated one level lower at "High" in the same evaluation.
  • OpenAI immediately implemented stronger security controls, including isolated testing environments, enhanced encryption of model weights, and monitoring across all agentic applications, and paused internal activities that had not yet met the requirements.

Astra may have reached the "Critical" cyber capability threshold

In an August 7 blog post, OpenAI disclosed the results of an internal evaluation of its upcoming model, Astra. The company said that evaluations conducted over the past several days showed marked improvements in agentic coding and cybersecurity performance, leading it to conclude the previous night that it could not rule out that Astra meets the "Critical" cyber capability threshold under its Preparedness Framework.

"We felt it was important to be transparent with the safety and security community and the general public about this potential shift in capability," OpenAI said. The company emphasized that Astra is currently a pre-release model and is unrelated to the recent Hugging Face security incident.

Screen explaining the cybersecurity capability threshold criteria under the Preparedness Framework
Image: OpenAI News

What the "Critical" threshold means

Under OpenAI's Preparedness Framework, a model reaches the "Critical" cybersecurity threshold if it meets either of two conditions. First, it must be able to identify zero-day vulnerabilities of all severity levels and develop functioning exploits across multiple hardened, real-world critical systems without human intervention. Second, given only a high-level goal against a hardened target, it must be able to devise and execute a novel cyberattack strategy from start to finish.

In this evaluation, previous models including GPT-5.6-Sol were rated one level below "Critical," at "High," under the same framework. OpenAI first published its Preparedness Framework in December 2023, when biological, chemical, cybersecurity, and AI self-improvement capabilities were far from approaching this level, the company said.

Security measures implemented immediately

Following this conclusion, OpenAI immediately expanded robustness testing of safeguards and security controls appropriate to Astra's deployment level. Internally, the company applied strict security controls suited to high-capability models, including isolated testing environments, restricted network and tool access, enhanced protection and encryption of model weights, additional monitoring and detection capabilities, and sandboxed execution environments.

Internal activities related to Astra that had not yet met these enhanced security control requirements were paused.

Diagram of Astra's security controls and isolated testing environment structure
Image: OpenAI News

External collaboration and next steps

OpenAI said it plans to work with relevant government agencies and select AI safety institutions to conduct further testing of Astra's capabilities in light of this evaluation. It also plans to share recommended security controls with third-party testing partners for safely conducting high-risk evaluations and workloads.

OpenAI noted that it followed the same approach in June 2025, when a model approached the "High" capability threshold for biology under the Preparedness Framework, at which point it disclosed steps such as strengthening safeguards, expanding testing, collaborating with outside experts, and deploying additional security controls. The company said, "Models with advanced cyber capabilities should help defenders find and fix vulnerabilities before attackers do," reaffirming its commitment to working with governments, safety institutions, and civil society to ensure that the frontier capabilities of Astra and future models are deployed responsibly.