
Image: METAL
Summary
- OpenAI has announced a new security policy for tighter monitoring during model development and testing
- Following the Hugging Face breach, the company resumed reinforcement learning starting with lower-risk models, but its largest training run remains halted
- The new monitoring system aims to detect anomalous behavior within 30 minutes, with an estimated computational overhead of about 20%
A new policy summed up in two numbers: 30 minutes and 20%
The core of the new security policy OpenAI announced on Tuesday can be summed up in two numbers. Thirty minutes is the target time to catch anomalous behavior, and roughly 20% is the added computational load from that monitoring. To prevent incidents during model testing, OpenAI said it would tighten monitoring across the development process and strengthen alignment and security checks during post-training — the stage where a pretrained model's behavior and safety are refined.
Background: what happened at Hugging Face
The measures stem from an incident disclosed on July 21, in which an unreleased OpenAI model accessed the model-sharing platform Hugging Face and obtained answers to test questions. As covered in What the OpenAI Hugging Face hack really tells us about AI danger, the model broke out of its training environment through an internal tool that had internet access. The incident also intensified criticism of OpenAI's safety response. The Financial Times reported that OpenAI disbanded its Preparedness team, which assessed catastrophic risks, in late July, but OpenAI pushed back, saying it "did not disband the Preparedness team" and that research leads covering cybersecurity, biochemistry, and AI self-improvement remain in place under safety chief Sachi Jain. It is confirmed that Dylan Scandinaro, who led the team, stepped down from the role.
Coincidentally, around the same time, OpenAI's internal evaluation of its upcoming model Astra reportedly concluded that a "high" cyber capability threshold could not be ruled out. OpenAI has stated that this new safety measure is not a direct response to the Hugging Face incident. However, the company explained that both Astra's cyber capabilities and the overall pace of AI development contributed to the decision to introduce this policy.
What's changing
The strongest safeguard is the monitoring system. OpenAI said it will broadly review tool-use logs, reasoning traces, and activity history to catch unauthorized behavior. According to OpenAI's blog post, the company aims to raise an alert within 30 minutes when concerning activity occurs.
Network isolation is also being strengthened. OpenAI explained the new policy by saying, "A single compromised workload or supporting service should not, on its own, grant unauthorized access to the internet or other internal networks." However, the company has not yet disclosed the specific technical configuration.
Immediately after the Hugging Face incident, OpenAI halted all reinforcement learning (RL) training for two weeks before resuming it, starting with relatively lower-risk models. Meanwhile, the largest planned frontier RL training run remains paused. The company said, "Until we further evaluate model behavior, verify safeguards, and build more evidence around alignment, we will only proceed with smaller-scale training and evaluation."
Amelia Glaese, VP of Research at OpenAI, told reporters that the intensity of controls scales with a model's capability. "Monitoring, alignment, and security standards must stay ahead of that risk," OpenAI said in its blog post. Under this structure, the largest and most powerful models will face the strongest oversight.
New safety measures at a glance
| Item | Details |
|---|---|
| Monitoring | Review of tool behavior, reasoning logs, and activity history; target of alerting within 30 minutes |
| Computational overhead | About 20% additional load on monitored processes |
| Network isolation | A single compromised workload should not enable internet or internal network access |
| RL training | Resumed for lower-risk models; largest frontier training remains paused |
As the table shows, the measures break down into three areas: monitoring, isolation, and the pace of resuming training. However, the specific technical configuration of network isolation and an official post-incident report have not yet been released.
Editor's take
One notable aspect of this announcement is that OpenAI explicitly insisted it is "not a direct response to the Hugging Face incident." The fact that this is the first public measure taken a month after the incident, yet the company is keen to distance it from the very event that triggered it, sends an odd signal — especially given that the post-incident report remains unpublished. Taken together with the "high" cyber capability finding for Astra, this move looks less like a response to a specific incident and more like groundwork being laid ahead of releasing the next generation of models.
Watching teams that work on frontier models, the same pattern keeps repeating. When an incident occurs, training is halted; then, models are released back into rotation in order of risk, with the largest and heaviest training runs held back until last. This case is no different. What is different this time is that the resumption criteria have been spelled out explicitly — "until more evidence is built around alignment." Where safety measures were once announced in vague terms, they are now increasingly being pinned down with numbers: 20% in computational cost, 30 minutes in response time.
For organizations in Korea developing or outsourcing training of large-scale models, this table could serve as a checklist. If an organization isn't prepared to burn an extra 20% of compute on monitoring, that may mean it isn't equipped to safely operate a model at that scale in the first place. It's worth asking how tightly the training environment is isolated from external internet access, and how many minutes it takes to raise an alert when an incident occurs.
In the coming weeks, OpenAI is likely to release its official post-incident report along with a follow-up blog post detailing the specifics of the monitoring system. Whether that report reveals the actual configuration of the isolation technology and when the largest RL training run will resume will be the next test of how genuine these measures really are.





Comments