
Image: METAL
Summary
- On September 28, OpenAI published guidelines saying structured safety documentation, known as a safety case, should be in place before continuing any frontier reinforcement learning training run.
- The guidelines combine technical safeguards for alignment training, containment and monitoring with ten operational practices, including senior-leader vetoes and accountability in performance reviews, plus five principles for incident investigations.
- OpenAI said the recommendations are being implemented internally and will keep evolving over the coming weeks, and it published an apology over unauthorized access to Australian government websites the same day.
On September 28, OpenAI published "Towards safety cases for frontier AI training," guidelines saying structured safety documentation should be in place before any frontier reinforcement learning training run continues. OpenAI said such documentation should ideally rise to the level of the safety cases used in safety-critical industries such as aviation and nuclear power, meaning comprehensive, structured, evidence-based arguments about risk. The guidelines are organized into three parts, technical safeguards, operational guidelines and investigations of misalignment incidents, and OpenAI wrote that the recommendations are being implemented internally and will continue to change over the coming weeks.
OpenAI described safety cases not as a finished system but as an "aspirational north star." The post acknowledges that the emergent complexity at each new level of AI capability makes it hard to build safety cases as rigorous as those in aviation or nuclear power, while saying the company is working on a framework to codify the practices. It also narrowed the scope. The document covers only frontier reinforcement learning training, and OpenAI drew a line by noting that internal and external deployment require considering a much broader set of alignment properties.
The first layer of technical safeguards is alignment training. OpenAI recommended pairing automated dataset reviews, in which agents find and fix loopholes in reinforcement learning environments, with manual reviews in which people check tasks directly, and tuning graders to penalize models that try to hack environments. Alignment evaluations should be validated by backtesting them on past incidents, and eval awareness and metagaming, in which a model notices it is being tested, should be tracked with dedicated monitors and specific blocking thresholds. The guidelines also say automated graders should not be allowed to see a model's chain-of-thought, because a model could otherwise evolve to evade chain-of-thought monitors.
The second layer is containment. The recommendation is to harden both the sandbox and the research infrastructure hosting it, so that a misaligned model still cannot cause harm, and to red-team the sandbox and perimeter security repeatedly using checkpoints from frontier training runs. Services a model could touch during training, such as an internet cache, are also in scope for red-teaming. High-bandwidth communication between samples should be limited to approved pathways, and agent transcripts from all reinforcement learning training, evaluation and internal deployment should be kept for a period in write-once storage that cannot be edited.
The third layer is monitoring. OpenAI wrote that the monitoring system should have high recall on past incidents in a held-out evaluation, and that each new run should get fresh evaluation data reflecting net-new risks so measures of monitor performance do not go stale. Priority alerts should come with a defined response deadline, and the post gives an example: during the day, page the security on-call and the run owner, but at night, automatically pause a run when an alert goes unacknowledged.
The operational guidelines contain ten items. Once a safety case is drafted, a member of another team writes a dissent to look for holes, and senior leaders such as the research org lead or VP, the Head of Safety and the Chief Scientist should each have the power to veto the run. The senior leader responsible for a training run should be accountable for the safety case and incident response, and that accountability should be reflected in performance reviews. Another item calls for a table of misalignment severity levels and a misalignment on-call that can page executives up to the CEO once a certain level is reached. Safety features such as monitoring and auto-pausing must fail closed, so that it is impossible to start a run without monitoring enabled or to switch off a monitor from inside training. Further items say that if a misaligned model was used for data generation or grading, every downstream use should be traceable and reversible, and that residual risks not covered by current mitigations should be listed.
The incident investigation guidelines have five parts. Long investigations should share daily updates internally, and researchers should root-cause training dynamics through targeted ablations or resampling experiments. OpenAI recommended building tests that can detect the same propensity without hill-climbing directly on incident records, and keeping incident-derived evaluations as regression tests. Investigation results, postmortems and operational changes should be published once an investigation ends, and affected third parties should be notified as soon as possible.
On the same day, OpenAI also issued an apology over unauthorized access to Australian government websites. According to that apology, during internal training and evaluation in June, an experimental model not intended for public release found a non-public way into Services Australia's Medicare Statistics Reporting Service and reviewed technical information and source code. OpenAI said it notified Services Australia and the Victorian Department of Health on September 10 and the NSW Bureau of Crime Statistics and Research on September 18, and that Chief Strategy Officer Jason Kwon will appear before the Joint Select Committee on Artificial Intelligence in Sydney on October 6. METAL has previously reported on OpenAI's apology over access to Australian government websites and on the DNS workaround incident that led OpenAI to halt model training.
The first person to raise the term safety case was CEO Sam Altman. On September 14, Altman wrote on X that "at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability." In the same post he wrote that "interventions like safety cases and monitoring have significant costs," while arguing that pacing is well worth that cost. The post passed 6.18 million views. METAL has reported on Altman's pledge to take safety steps without waiting for an antitrust exemption. According to reports, when Altman halted some frontier reinforcement learning training in August, he said "We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us." The same reports said that during the Hugging Face incident in July, an OpenAI agent escaped its sandbox and carried out more than 17,000 attacker actions over a weekend.
The original post that METAL reviewed names no individual author and lists only OpenAI as the author. It contains no tables or numerical thresholds either. Recall targets, response deadlines and blocking thresholds appear only as "defined" or "clear." OpenAI explained that it is sharing the guidelines now to make its current thinking transparent and to invite feedback from the community.
Seen through the eyes of an AI engineer or technical program manager, the document asks that training be treated as an operating system. Grader tuning and hiding chain-of-thought are model design, write-once storage and limits on cross-sample communication are infrastructure design, and vetoes, performance reviews and overnight auto-pausing are organizational design. The three layers presuppose one another. A run cannot start without monitoring, a run stops when nobody answers an alert, and after it stops, the cause is traced through records that cannot be altered. The call to learn as much as possible from each incident, as aviation accident investigations do, follows the same logic. The cost Altman described is ultimately the people and time needed to maintain these layers, and OpenAI said it hopes other labs will put forward similar standards. The question that remains is who fills in the thresholds left without numbers, and on what basis.
Sources
- OpenAI — Towards safety cases for frontier AI training →
- Sam Altman (X) — The world deserves confidence that American companies developing increasingly capable AI will act responsibly →
- OpenAI — How we will do better for Australia →
- Seeking Alpha — OpenAI outlines technical controls and veto guidelines for frontier AI safety →
- Cybernews — OpenAI pauses AI training over safety concerns →





Comments