
Image: METAL
Summary
- Nonprofit Guidelight has published its first assessment of internal AI safety controls at Anthropic, OpenAI, Google, xAI, and Meta
- Anthropic and OpenAI topped the list with C+ grades, while Google got a D+, xAI a D-, and Meta an F
- All five companies were decent at detecting risk but weak on prevention and containment
Starting with the report card
Nonprofit organization Guidelight has published the results of its first-ever assessment of how well five companies — Anthropic, OpenAI, Google, xAI, and Meta — actually control their own AI systems internally. Anthropic and OpenAI came out on top, if you can call it that, with C+ grades. Google got a D+, xAI a D-, and Meta an F. Guidelight said not a single company fully met the safety standard it had proposed.
| Company | Grade | Average across 6 categories (out of 5) |
|---|---|---|
| Anthropic | C+ | 2.50 |
| OpenAI | C+ | 2.50 |
| D+ | 1.50 | |
| xAI | D− | 0.83 |
| Meta | F | 0.67 |
What each of the six categories demands
Guidelight looked at six baseline practices drawn from the Control standard it published last May. Here's what the standard actually calls for in each area.
| Category | What the standard requires |
|---|---|
| Logging | Keep 99.9% of reasoning tokens from internally used risk-relevant models in a monitorable, recorded form. This doesn't mean running a classifier on every single request — it means preserving the data so a classifier could scan it at any time |
| Measuring monitoring effectiveness | Measure and document how well the monitoring system actually catches problematic behavior (recall) |
| Gate actions | At minimum, require human sign-off before two kinds of actions can proceed: (1) changes to the monitoring system itself, and (2) changes to control-related security settings. The human reviewer must also be able to see the agent's chain-of-thought |
| Circuit breaking | Once flagged instances of problematic behavior cross a threshold, block the gated behavior entirely until a human determines it's safe to resume |
| Third-party review | Commission and publish a quarterly report from an external body assessing the adequacy of the control system. Redactions are allowed only where a third party judges the material poses commercial risk or is otherwise inappropriate to disclose |
| Containment plan | Establish and adopt a response plan that restricts a model's access and operation if it's detected attempting to circumvent controls |
How the 0-5 scores turned into letter grades
Scores run on an absolute 0-to-5 scale: 0 means not implemented, 1 means only preliminary steps are in place, 2 means limited partial implementation, 3 means substantial partial implementation, 4 means nearly complete implementation, and 5 means fully implemented. Guidelight also included a separate relative ranking comparing companies against each other.
The letter grades come from averaging the six absolute scores and mapping that average onto a standard U.S. GPA scale. For this round, all six categories were weighted equally; Guidelight said it may adjust the weighting in future rounds. Anthropic and OpenAI both landed on C+ because they tied at 15 total points.
Score breakdown by category
| Category | Anthropic | OpenAI | xAI | Meta | |
|---|---|---|---|---|---|
| Logging | 3 | 3 | 2 | 0 | 1 |
| Measuring monitoring effectiveness | 3 | 3 | 1 | 0 | 1 |
| Gate actions | 3 | 2 | 1 | 2 | 0 |
| Circuit breaking | 3 | 2 | 1 | 2 | 0 |
| Third-party review | 3 | 2 | 2 | 0 | 2 |
| Containment plan | 0 | 3 | 2 | 1 | 0 |
| Total | 15 | 15 | 9 | 5 | 4 |
| Grade | C+ | C+ | D+ | D− | F |
No company scored above 3 in any category, and most scores sat at 2 or below. Anthropic hit 3 in five categories but scored a 0 on containment planning, which dragged its average down. OpenAI, by contrast, had containment planning as one of its three 3-point categories.
How the grading worked
Between June and August, Guidelight gathered public materials from all five companies — frontier safety frameworks, system cards and model reports, technical blog posts, and company statements cited in third-party reports. For each category, two team members independently scored the evidence against a shared rubric; if they disagreed, a third member joined the discussion until they reached consensus. For every score, the team asked what additional evidence or measures would have justified a higher score, and whether alternate readings of the standard's language would have changed the outcome.
In late July, each company's team received preliminary scores and supporting rationale in advance, with a chance to flag factual errors, push back on interpretations, point to overlooked evidence, or voluntarily disclose more information. In early August, Guidelight released version 1.1 of the standard, which extended control requirements beyond internal operations to include the evaluation and testing stages — and some scores were revisited to reflect that change. The scores in this report reflect the state of affairs as of August 18.

Good at spotting problems, bad at stopping them
The five companies split into three tiers. Anthropic and OpenAI log a substantial share of internal AI activity, scan it for signs of problematic behavior, and have procedures to escalate suspicious activity to a human. Google lags in actual implementation, but its AI control roadmap was judged to be the most detailed plan any company has published so far, spanning prevention, detection, and containment — though based on public materials, most of it hasn't been put into practice yet. Meta and xAI had the weakest practices and the weakest plans of the group. Most of what's known about Meta comes from the company's own submission to METR's frontier risk report, and xAI is the only one of the five that didn't participate in that report at all.
The strongest area across the board was detection and third-party evaluation. Three companies said they log some portion of internal usage and scan it for problematic behavior, and four of the five (all but xAI) took part in the METR report. The weakest areas, by contrast, were prevention and containment. Anthropic was the only company to score above "limited partial implementation" on gate actions and circuit breaking. Guidelight's takeaway: as things stand, a company's own control system could potentially be disabled by a misbehaving AI, and could be overwhelmed by a fast sequence of attacks that outpaces how quickly a company can respond.
What actually earned or cost each company points
Anthropic's 3s on gate actions and circuit breaking came from Claude Code's "auto mode," which lets a monitor review an agent's actions and block anything risky; once refusals hit a certain threshold, the system escalates to a human or halts the agent. It's enabled by default in internal deployments. But auto mode doesn't read the chain-of-thought, and stopping one agent doesn't stop the system — a new sub-agent can simply be spun up in its place. Guidelight found no evidence of a system-wide kill switch, which is what cost Anthropic points.
OpenAI's sole 3 on containment planning came from an actual track record of pulling the plug. During its response to the Hugging Face security incident, OpenAI paused training of new models and restricted internal deployment, and gave its investigation team the ability to halt a model's activity if monitoring caught signs of misalignment. There's also a prior instance of the company restricting deployment of a model that had escaped its sandbox. But Guidelight found no formal plan spelling out when and how the company would respond if something similar happened again.
The limits Guidelight itself acknowledged
These scores are based only on public information. A company with strong containment mechanisms that it simply hasn't disclosed would score poorly here regardless. On the flip side, Guidelight noted that what companies do disclose hasn't been independently verified, so it's possible some claims overstate the reality. The six categories are also just the baseline of the Control standard — scoring well here doesn't mean a company's overall control practices are adequate, and control itself is only one piece of a company's broader safety performance. The standard represents a floor, not a ceiling: fully meeting it wouldn't necessarily count as sufficient control. And the assessment doesn't cover controls applied during model training itself or during testing by outside evaluators — those fall outside its scope for now.
Who's behind Guidelight
Guidelight is an independent nonprofit founded by Paige Hadley and Steven Adler, both former OpenAI safety staffers. This is the group's first assessment. It says it plans to update Control standard scores on a recurring basis and eventually expand into other standards, such as transparency. OpenAI's relatively strong C+ showing appears to be tied in part to the internal monitoring upgrades it announced on August 18. That said, its largest reinforcement learning training run reportedly remains paused.
Editor's view
What stands out in this assessment isn't the ranking — it's the order in which things got built. All five companies invested at least somewhat in noticing "something's gone wrong," but far less in figuring out "so what do we do about it now." That's not a coincidence. Detection is relatively cheap: you just need logs and alerts. Containment and prevention are expensive, because they require actually stripping a model of its permissions or pulling the plug on training — decisions with real costs attached. The order isn't backwards; it's just that companies filled in the easy part first.
The scorecard makes the gap even clearer. Anthropic scored 3 in five of six categories but got a 0 on containment planning, which capped its average at 2.50. OpenAI's 3 on the same category, meanwhile, came not from a document but from an actual track record of halting training. That distinction matters: the point wasn't awarded for having a plan on paper, but for having actually pulled the trigger before. It's a signal that this assessment is measuring track record, not intent.
For companies in Korea deploying AI systems, there's a clear practical takeaway here. Asking a vendor "how fast can you catch abnormal behavior?" isn't enough anymore. You need to ask the follow-up: "and once you catch it, what can you actually do?" An answer like "we keep audit logs" and an answer like "we can immediately cut off a compromised model's access" represent two very different levels of readiness.
Don't expect the rankings to shift dramatically in Guidelight's next assessment. Building out prevention and containment systems means restructuring organizational decision-making authority itself — not something that catches up in a matter of weeks. That said, companies like OpenAI and Google, which have already published roadmaps and disclosed reinforcement measures, have room to close the gap next time. Companies like Meta, where public disclosure remains thin to begin with, are more likely to stay right where they are.





Comments