METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

OpenAI publishes principles for third-party safety assessments

A single document sets out four priority areas for deeper outside assessment and seven principles the work should follow. Claims are pre-registered before assessment, and assessors can note in their reports when the company has requested redactions.

OpenAI publishes principles for third-party safety assessments

Image: METAL

Summary

  • On September 22, OpenAI published four priority areas and seven principles for third-party safety assessments.
  • The priority areas are safety cases, critical safeguards, Preparedness Framework capability evaluations, and investigations of critical misalignment incidents.
  • Claims are pre-registered, conflicts of interest are disclosed, and assessors keep editorial independence and the right to flag redactions.

OpenAI on September 22 published principles for how outside organizations should verify the safety of its models. The document lays out four areas for deeper third-party assessment and seven principles the assessments should follow, and commits to giving assessors deep levels of access across training, evaluation, and deployment. That access, the company wrote, "should enable assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards."

The starting point is how responsibility is divided. OpenAI wrote that frontier AI labs carry an immense responsibility to train, evaluate, and deploy models safely, and that third-party assessments are a critical way to balance that responsibility and keep labs accountable to clear, independently supported safety claims. The scope is limited to technical safety assessments by independent organizations in the private and non-profit sector; testing and evaluation work with governments is set aside, since different roles and responsibilities may call for different approaches.

The document begins with definitions. A safety claim is a specific assertion about a model or system's capabilities, behavior, or safeguards that can be assessed against evidence, and it should state the risks and conditions it covers along with its assumptions and limitations. A safety case is a structured, evidence-backed argument that a system's risks are adequately managed for a specified activity such as training, evaluation, or deployment. It connects individual safety claims to their evidence and makes explicit the assumptions, uncertainties, and remaining risks that could affect its conclusions.

The first area for deeper assessment is the safety case itself. It spans training, evaluation, internal deployment, and external deployment, and the company wrote that it calls for expertise in alignment, control methods such as monitoring, cybersecurity, biological and chemical misuse, and red teaming. Assessors are to examine whether the evidence behind safety cases is substantiated, whether the conditions of the safety case were actually followed, and whether training contains incentives that could reward deception, reward hacking, destructive actions, or circumventing restrictions.

The second is critical safeguards across internal and external deployments. OpenAI's safeguards include model-level safeguards, enforcement safeguards, security safeguards, and misalignment monitors. Assessors would use "grey box" access to test whether safeguards hold up against jailbreaks, and examine how agents interact with access controls, sandboxing, and detection and response systems under authorized, realistic operating conditions. How reliable chain-of-thought monitoring remains as evidence as model capabilities improve is also on the list of questions.

The third is capability evaluations. The target is evaluations covering the risk categories in OpenAI's Preparedness Framework, namely chemical and biological risks, cybersecurity, and AI self-improvement, along with alignment evaluations for severe misalignment. When models consistently hit top scores and evaluations saturate, outside assessors are meant to check whether new tests meaningfully measure more advanced capabilities and whether thresholds are set correctly.

The fourth is independent investigation of critical misalignment incidents. The document cites models acting without authorization or evading oversight, and says that in select circumstances, as with the OpenAI Hugging Face incident, bringing in an independent third party can be beneficial. METAL has reported that the UN AI scientific panel described that incident as an early warning of one pathway to loss of control, and that the independent non-profit METR counted agent activity during the period it reviewed; the new document turns such investigations from a one-off into a standing area. Investigators are expected to bring cyber forensics and alignment expertise, large-scale chain-of-thought analysis, and enough staff to finish in a timely manner.

The seven principles read like contract clauses. Assessments begin with a mutually agreed scope, and the safety claims to be assessed are pre-registered before work starts. The parties must also say whether a claim is one the company put forward or one the assessor set independently. Access is proportionate to the agreed claims within legal, security, and IP constraints, and where direct access is impractical, assessors can work through a designated company representative or privacy-preserving indirect mechanisms. Assessors must explain their methods, criteria, and uncertainties, and reports must separate direct findings from interpretation.

From a lawyer's perspective, the weightiest provisions are independence and publication. Assessors must disclose conflicts of interest, including financial incentives, relationships with developers, and prior involvement in the work being assessed, and safeguards such as recusal or exclusion periods can keep commercial pressure and compensation from shaping findings. The company gets a reasonable period to remediate issues before publication and can request redactions of sensitive information, but assessors keep editorial independence and can note where substantive redactions were made and how they affected the assessment. Where full public disclosure is not possible, confidential reporting to oversight bodies such as boards remains an option.

Assessments will run from several weeks to several months. According to the document, OpenAI expects to support multiple assessments in parallel, and the work is generally launch-agnostic, digging into particular safety claims over time. The company said it has already given outside assessors information about its technical safeguards, visible chain of thought, and unprecedented levels of confidential data and internal deployment access for incident response and monitor red teaming.

The document builds on a commitment made about six weeks earlier. In its August 8 global policy newsletter, OpenAI first set out six principles that distinguish audits from third-party assessments. Audits examine whether an organization follows established requirements and processes, while third-party assessments test whether evidence supports specific safety claims. In that piece the company wrote that "requiring companies to undergo review is not enough," and said the system should guard against "assessor shopping" aimed at softening adverse findings. The same piece notes that OpenAI has worked with METR, SecureBio, Apollo Research, and Irregular to evaluate risks including long-horizon autonomy, biological threats, cybersecurity, scheming, deception, and oversight subversion.

The September 22 document METAL reviewed does not name any new assessment partners, contract sizes, or timelines. Its final paragraph ends by saying the company is in conversation with multiple third parties about proposals aligned with the priority areas, and Lama Ahmad is listed as the author. The company also introduced the four priority areas and principles on its official X account, where the post had passed 499,000 views when METAL checked.

Similar moves are under way across the industry. METAL has reported that Anthropic signed an independent evaluation partnership with Accenture as a first step toward its pledge to embed evaluators inside the company. OpenAI's document differs in that it spells out what is assessed and under which conditions before who does the assessing, and the company said it intends to turn those standards into shared international standards through both future laws and private governance institutions.

In legal terms, the document's weight lies less in its promises than in its procedures. Pre-registered claims, conflict-of-interest disclosure, and flagged redactions are all mechanisms that make assessment results contestable later, and they are the provisions regulators are likely to consult first when they require third-party verification. OpenAI wrote that "no one third party can or should comprehensively cover urgent frontier safety questions." Whether the principles work in practice will show first in what the initial assessment reports redact and keep, and whether assessors state that themselves.

Comments