One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

OpenAI's Upcoming Model Astra Flags Potential 'Critical' Cyber Capability

OpenAI disclosed on August 7, 2026 that internal evaluations of its upcoming model, Astra, could not rule out a "Critical" cyber capability rating under its Preparedness Framework. The prior model, GPT-5.6-Sol, had been rated "High" on the same scale.

이미지: OpenAI News

Summary

  • OpenAI disclosed that its upcoming model, Astra, came close to the "Critical" cyber capability threshold in internal evaluations.
  • OpenAI is immediately applying enhanced security controls to Astra's continued development, including isolated testing environments, restricted network access, and encryption of model weights.
  • The company plans to work with relevant government agencies and AI safety institutions to further verify the model's capabilities.

Astra's Potential 'Critical' Cyber Capability Disclosed for the First Time

OpenAI is a company that has pursued AI research and development under a mission to ensure that artificial general intelligence (AGI) benefits all of humanity. It operates under a structure in which the nonprofit OpenAI Foundation governs and oversees the for-profit OpenAI Group, with both entities working together to advance the mission.

On August 7, 2026, OpenAI disclosed the results of internal evaluations conducted over recent days on Astra, a model that has not yet been released. According to the evaluation, Astra showed notable performance improvements in agentic coding and cybersecurity, and after incorporating expert review, the company concluded that it could not rule out a "Critical" cyber capability rating under its Preparedness Framework.

Diagram explaining the criteria for cybersecurity capability ratings under the Preparedness Framework
Image: OpenAI News

What Is the Preparedness Framework

OpenAI first published its Preparedness Framework in December 2023. The framework is designed to track the level of model capability reached in key areas — including biological, chemical, cybersecurity, and AI self-improvement domains — and to predefine the actions the company must take at each capability inflection point.

Under the framework, the "Critical" cybersecurity threshold has specific criteria. A model reaches this rating if it can identify and develop functional zero-day exploits of all severity levels across multiple real-world critical systems without human intervention, or if it can independently devise and execute an entirely novel cyberattack strategy against a hardened target given only a high-level goal. The prior model, GPT-5.6-Sol, was rated "High" under the same criteria, falling short of the Critical threshold. OpenAI noted that Astra is still undergoing benchmarking and evaluation, and that this matter is unrelated to the Hugging Face security incident.

Immediate Actions Taken by OpenAI

Immediately after reaching this conclusion, OpenAI expanded robustness testing of its safeguards and security controls to ensure deployment matches the model's cyber capability level. Internally, the company is pursuing several additional measures in parallel.

First, it has introduced enhanced security controls appropriate for high-capability models, including isolated testing environments, restricted network and tool access, strengthened protection and encryption of model weights, additional monitoring and detection capabilities, and sandboxed execution environments. Internal activities related to Astra that do not yet meet these enhanced requirements have been temporarily suspended. OpenAI has also introduced universal monitoring for risky behavior and misalignment across all agentic applications of Astra, including training and evaluation. This monitoring system evaluates the model's chain of thought in real time, and if high-risk activity is detected, it triggers a security response to review and halt the activity.

Screen listing Astra's security control measures
Image: OpenAI News

Externally, OpenAI plans to work with relevant government agencies and select AI safety institutions so they can independently verify the model's capabilities, and it intends to provide third-party testing partners with recommended security control guidelines for safely conducting high-risk evaluations and tasks.

Prior Precedent and Context for the Disclosure

OpenAI stated that this is not the first time it has responded to a capability transition. In June 2025, when one of its models approached the "High" capability threshold for biological risk under the Preparedness Framework, the company similarly disclosed procedures including strengthened safeguards, expanded testing, collaboration with outside experts, and additional security controls. OpenAI said it is applying the same principles to this cyber capability matter.

Through this disclosure, the company emphasized public transparency. OpenAI stated that models with advanced cyber capabilities should help defenders discover and fix vulnerabilities before attackers can exploit them, and that it will work with governments, safety institutions, and civil society to ensure that Astra and future models' frontier capabilities are deployed responsibly and in ways that benefit all of humanity. No specific plans, such as a deployment timeline or commercialization schedule for the model, have been disclosed at this time.