One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

OpenAI's upcoming model Astra approaches critical cyber capability

On August 7, 2026, OpenAI disclosed that internal evaluations of its upcoming model Astra could not rule out "Critical" cybersecurity capability under its Preparedness Framework. Previous models, including GPT-5.6-Sol, had stayed at "High" on the same scale.

이미지: AI 생성 — METAL LAB

Summary

  • OpenAI announced that cybersecurity evaluations of its upcoming model Astra could not rule out a "Critical" rating under its Preparedness Framework.
  • Astra was assessed as approaching the level at which a model could develop zero-day exploits against hardened real-world systems without human intervention, or devise and execute end-to-end cyberattack strategies from only a high-level goal.
  • OpenAI immediately tightened internal security controls—building isolated testing environments, restricting network and tool access, and strengthening protection of model weights—and paused some internal work on Astra.

What changed

On August 7, 2026, OpenAI disclosed the results of internal evaluations conducted over the past few days on its upcoming model, Astra. The evaluations showed meaningful performance gains in agentic coding and cybersecurity, and after incorporating outside expert assessments, the company concluded the previous night that it could not rule out "Critical" cyber capability under its own Preparedness Framework. OpenAI said transparency was the reason for disclosing this potential capability shift, and it aimed the disclosure at both the safety and security community and the general public.

OpenAI first published its Preparedness Framework in December 2023. At that time, biological, chemical, cybersecurity, and AI self-improvement capabilities were nowhere near this threshold. The framework was designed to track capability progress and define what actions the company should take once those capabilities materialize. Previous models, including GPT-5.6-Sol, were classified as "High" rather than "Critical" in frontier cyber capability evaluations.

Diagram explaining the Preparedness Framework's cybersecurity capability rating system
Image: OpenAI

What qualifies for a "Critical" rating

Under the Preparedness Framework, a model qualifies for the Critical cybersecurity rating if it meets either of two conditions. First, it can identify and develop functional zero-day exploits of any severity level against multiple hardened, real-world core systems without human intervention. Second, given only a high-level goal, it can independently devise and execute an end-to-end, novel cyberattack strategy against hardened targets.

OpenAI said benchmarking and evaluation of Astra are still ongoing, but preliminary results already showed performance strong enough that a Critical capability rating could not be excluded. The company noted, however, that Astra is an upcoming model and is unrelated to the previously reported Hugging Face breach.

Immediate steps taken by OpenAI

Following the evaluation results, OpenAI immediately implemented a series of internal security measures. These include building isolated testing environments, restricting network and tool access, strengthening protection and encryption of model weights, adding monitoring and detection capabilities, and applying sandboxed execution environments. The company also paused internal work related to Astra that had not yet met the enhanced security control requirements.

In addition, OpenAI introduced general-purpose monitoring for risky behavior and misalignment across all of Astra's agentic applications, including training and evaluation. The monitor evaluates the model's chain of thought, and when high-risk activity is detected, it triggers a security review that halts the activity. Externally, OpenAI plans to work with relevant government agencies and select AI safety institutions to verify the model's capabilities, and it intends to share recommended security controls with third-party testing partners to help them safely conduct high-risk evaluations and workloads.

Flowchart of Astra's agentic monitoring and security controls
Image: OpenAI

Precedent and what comes next

This is not the first time the Preparedness Framework has guided the company's response to an actual capability transition. In June 2025, when OpenAI's models approached the "High" biological capability rating under the framework, the company similarly announced strengthened safeguards, expanded testing, collaboration with outside experts, and additional security controls. The same principles are being applied to this cybersecurity Critical capability evaluation.

OpenAI said models with advanced cyber capabilities should help defenders find and fix vulnerabilities before attackers can exploit them. The company added that it would continue working with governments, safety institutions, and civil society to ensure that the frontier capabilities of Astra and future models are deployed in ways that benefit humanity as a whole. This announcement did not include specific commercialization plans, such as a deployment timeline or conditions for Astra's external release.