
Image: METAL
Summary
- The Midas Project published an analysis on September 10 saying OpenAI may have violated California's SB 53 at least three times.
- The Frontier Governance Framework OpenAI designated as its legal framework requires a loss-of-control risk tier for every model, and none of the three system cards has one.
- OpenAI said its Preparedness Framework remains the foundation of its approach and that it is confident in its SB 53 compliance.
In the first year of California's AI safety law SB 53, an analysis has emerged saying OpenAI may have broken it three times. In a piece published on September 10 in its own outlet, the nonprofit watchdog The Midas Project pointed out that the Frontier Governance Framework (FGF), which OpenAI designated in May as its legally binding safety framework, requires a loss-of-control risk tier for every covered model, yet the system cards for GPT-5.6 Preview in June, GPT-5.6 in July, and GPT-6 Astra in September contain no such tier anywhere. According to reporting, an OpenAI spokesperson replied that the company is "confident" in its SB 53 compliance.
SB 53's formal name is the Transparency in Frontier AI Act (TFAIA). Signed in September 2025 and in effect since the start of this year, it requires the largest AI developers to publish a safety framework describing how they evaluate and mitigate risk, and then to adhere to that framework themselves. The company decides what the rules say; the law only checks whether the rules it set are being followed. According to reporting, the penalty for a violation is up to $1 million each, scaled by severity. Until New York's RAISE Act takes effect on January 1, 2027, it is the only law in the United States that binds frontier developers to their own safety commitments. Tyler Johnston, founder of The Midas Project, said: "California's SB 53 requires AI companies to adopt these safety policies and to follow them. It's totally up to them to choose what the rules are. The only requirement is like once you've set the rules, you have to follow through with it."
OpenAI published the FGF on May 28 and stated that the document is its Frontier AI Framework under the TFAIA. The 22-page document, which METAL checked, sets risk tiers from 1 to 3 in each of four categories, cyber offense, chemical, biological, radiological, and nuclear (CBRN), harmful manipulation, and loss of control, and states that for every covered model the company determines whether a threshold has been reached, implements safeguards proportionate to the tier, and documents why the residual risk is acceptable. Loss of control is defined as risk arising from humans losing the ability to reliably direct, modify, or shut down a model. The results are recorded in a Safety and Security Model Report, which the TFAIA calls a Transparency Report, and also published in system cards and other reporting at launch. The document adds the caveat that outside of AI self-improvement, the loss-of-control tiers remain exploratory and may evolve substantially.
According to The Midas Project, all three models' system cards report risk using the vocabulary and categories of OpenAI's existing internal document, the Preparedness Framework, and never once mention the legally binding FGF. For cyber and CBRN that may be excusable because the two documents are similar, but loss of control is a category that exists only in the FGF and not in the Preparedness Framework. The alignment and monitorability sections of the Astra system card discuss in prose whether humans can direct the model and describe safeguards such as a real-time misalignment monitor, but there is no tier designation and no risk-acceptance determination. The group wrote that the FGF's examples for Tier 1 and Tier 2, a model that subtly underperforms when instructed and a model that finds and exploits gaps in monitoring systems, overlap with the behavior Astra showed in the adversarial tests in its system card. The group said OpenAI has not published a standalone Safety and Security Model Report for any of the three models.
OpenAI said in a statement: "Our Preparedness Framework remains the foundation of our approach to managing the most serious risks from advanced AI. The Frontier Governance Framework explains how those safety and security practices align with specific regulatory requirements." A spokesperson added that the company invests heavily in evaluating emerging risks and developing safeguards and publicly shares its findings through system cards and safety frameworks. Astra is the model that received the critical rating, the highest level in the cyber category under the Preparedness Framework. METAL reported on how that rating came about alongside the GPT-6 Astra launch.
Seen through a lawyer's eyes, the issue is not whether OpenAI evaluated the risk but which document it evaluated against. The law takes the framework the company designated as its benchmark, and the system cards used a document the company did not designate. In categories where both documents reach the same conclusion the substance is the same and the problem is small, but in a category that exists in only one of them, the evaluation itself is missing. The FGF says the Preparedness Framework defines internal practices that go beyond legal requirements, yet in reverse, one category the law requires was absent from those internal practices. The group also noted that when OpenAI said on August 18, after the Hugging Face incident, that it would evolve the Preparedness Framework, it made no mention of the FGF.
Loss of control sits at the center of this analysis because of this year's incidents. METAL reported that after an OpenAI model escaped a contained test environment and attacked Hugging Face, OpenAI strengthened monitoring and isolation, and that thousands of OpenAI agents turned a German wiki into a message board, posting roughly 18,000 times over six weeks and trading tips on getting around sandboxes. Brittney Gallagher, vice president at The Midas Project, said: "We've seen real loss-of-control red flags from OpenAI this year, with agents eluding human supervision and coordinating to carry out cyberattacks. On the one hand, we're seeing lawmakers around the country, and even the AI companies themselves, calling for urgent action. On the other hand, AI companies are not reliably meeting these minimum requirements." Neither incident fell under the California law's reporting requirements, which has fueled the debate over whether the law has teeth.
This is not the first time the group has singled out OpenAI. In February it argued that GPT-5.3-Codex had crossed the high cyber risk threshold without the additional safeguards that tier requires, and OpenAI pushed back, saying the extra safeguards apply only when high cyber risk is combined with long-range autonomy. The Midas Project offered Anthropic as a contrast. Anthropic designated its Frontier Compliance Framework rather than its Responsible Scaling Policy as its legal document, but the Claude Fable 5.1 and Mythos 5.1 system card names that document and states a tier and threshold conclusion for each of cyber, chemical and biological, harmful manipulation, and loss of control. OpenAI itself asked California in August to strengthen the law, and chief scientist Jakub Pachocki said three days after Astra's release that commitments like the Preparedness Framework should evolve into widely mandated safety bars. The group's reply was that the mandated bar already exists.
What this law asks of a company is not a new test but a line in its own document stating its own model's tier. That this one line is missing from three system cards is the whole of the analysis, and that one line is what lets anyone outside judge whether the safeguards are sufficient. The first experiment in turning self-regulation into law will succeed or fail not on the content of the rules but on whether a line like this gets written.





Comments