
이미지: The Decoder
Summary
- Nonprofit Guidelight has released its first assessment of the internal AI safeguards at Anthropic, OpenAI, Google, xAI, and Meta
- Anthropic and OpenAI topped the ranking with C+, followed by Google's D+, xAI's D-, and Meta's F
- All five companies are relatively good at detecting risk but weak on prevention and containment measures
- 평가 기관
- 가디라이트(Guidelight) — 전 오픈AI 안전 담당자 페이지 헤들리·스티븐 애들러 설립
- 평가 대상
- 앤스로픽, 오픈AI, 구글, xAI, 메타
- 등급
- 앤스로픽·오픈AI C+(2.50) / 구글 D+(1.50) / xAI D−(0.83) / 메타 F(0.67)
- 점검 항목
- 로깅, 모니터링 효과 측정, 게이트 액션, 서킷 브레이킹, 제3자 검토, 봉쇄 계획 6개
- 채점 척도
- 0~5점 절대 척도(0 미구현~5 완전 구현) + 회사 간 상대 평가 병기
- 등급 산정
- 여섯 항목 절대 점수의 단순 평균을 미국 GPA 환산표에 대입. 이번 회차는 항목 가중치 동일
- 평가 방식
- 공개 자료만 근거. 팀원 2명이 공통 기준표로 독립 채점 후 제3자 포함 합의, 7월 말 각 사에 예비 점수 회람
- 기준 시점
- 2026년 8월 18일 · 통제 표준 v1.1 기준
- 결과 패턴
- 탐지·제3자 평가는 상대적 강세, 예방(게이트·서킷 브레이킹)과 봉쇄는 전반적 취약
- 결론
- 어떤 회사도 어떤 항목에서도 3점을 넘지 못함
Starting with the report card
Nonprofit organization Guidelight has published the results of its first assessment of how well five companies — Anthropic, OpenAI, Google, xAI, and Meta — control their AI systems internally. Anthropic and OpenAI came out slightly ahead with C+, while Google received a D+, xAI a D-, and Meta an F. Guidelight said that in this assessment, not a single company fully met the safety standard it proposed.
| Company | Grade | Average across 6 items (out of 5) |
|---|---|---|
| Anthropic | C+ | 2.50 |
| OpenAI | C+ | 2.50 |
| D+ | 1.50 | |
| xAI | D− | 0.83 |
| Meta | F | 0.67 |

What each of the six items requires
What Guidelight evaluated were six basic practices drawn from the Control standard it published last May. The levels required by the original standard are as follows.
| Item | Level required by the standard |
|---|---|
| Logging | Record 99.9% of reasoning tokens from internally used risk-relevant models in a form that can be monitored. This doesn't mean running a classifier on every single request, but rather ensuring the data is retained so a classifier could scan it at any time |
| Measuring monitoring effectiveness | Measure and document how well the monitoring system actually catches problematic behavior (recall) |
| Gated actions | At minimum, two categories of action — ① changes to the monitoring system itself and ② changes to control-related security settings — must be approved by a monitor before execution, and the monitor must be able to view the agent's chain-of-thought |
| Circuit breaking | Once flagged problematic-behavior incidents accumulate past a threshold, gated actions should be blocked entirely until a human determines it is safe to resume |
| Third-party review | Have an external body produce a quarterly report assessing the adequacy of the control system and publish it. Redactions should be limited to portions a third party judges to pose commercial risk or to be inappropriate for disclosure |
| Containment plan | Establish and adopt a response plan that restricts a model's access and operation when attempts by that model to disable controls are detected |
How scores from 0-5 were converted into grades
Scores use an absolute 0-5 scale. 0 means not implemented, 1 means only preliminary steps are in place, 2 means limited partial implementation, 3 means substantial partial implementation, 4 means near-complete implementation, and 5 means full implementation. Guidelight also included a separate relative ranking comparing companies against each other.
The letter grades were derived by taking a simple average of the absolute scores across the six items and mapping that onto a U.S. GPA conversion table. In this round, all six items were weighted equally; Guidelight said future rounds may adjust the weighting of individual items. Anthropic and OpenAI both landed at C+ because their total scores were tied at 15 points.
Score breakdown by item
| Item | Anthropic | OpenAI | xAI | Meta | |
|---|---|---|---|---|---|
| Logging | 3 | 3 | 2 | 0 | 1 |
| Measuring monitoring effectiveness | 3 | 3 | 1 | 0 | 1 |
| Gated actions | 3 | 2 | 1 | 2 | 0 |
| Circuit breaking | 3 | 2 | 1 | 2 | 0 |
| Third-party review | 3 | 2 | 2 | 0 | 2 |
| Containment plan | 0 | 3 | 2 | 1 | 0 |
| Total | 15 | 15 | 9 | 5 | 4 |
| Grade | C+ | C+ | D+ | D− | F |
No company scored above 3 on any single item, and most scores were 2 or lower. Anthropic scored 3 on five of the six items but was dragged down by a 0 on containment planning, while OpenAI's containment planning score was one of its own 3-point items.
Methodology
Between June and August, Guidelight gathered publicly available materials from all five companies, including frontier safety frameworks, system cards and model reports, technical blog posts, and company statements quoted in third-party reports. For each item, two team members scored independently using a shared rubric, and when they disagreed, a third member joined the discussion until consensus was reached. For each score, the team revisited questions such as what additional evidence or measures would have raised the score, and whether alternate readings of the standard's wording would change the outcome.
In late July, Guidelight sent preliminary scores and supporting rationale to each company's staff in advance, giving them a chance to flag factual errors, contest interpretations, point out overlooked materials, or voluntarily disclose additional information. In early August, Guidelight released standard v1.1, which expanded the scope of control requirements beyond internal operations to include the companies' own evaluation and testing stages, and it revisited some scores affected by that change. The cutoff date for this report card is August 18.
Good at detection, weak at prevention
The five companies split into three tiers. Anthropic and OpenAI log a substantial portion of internal AI activity, scan it for signs of problematic behavior, and have procedures to escalate suspicious behavior to humans. Google lags in implementation, but its AI control roadmap was judged to be the most concrete planning document released by any company so far, covering prevention, detection, and containment — though based on public materials, most of it has yet to be implemented. Meta and xAI had the weakest practices and plans of all. Most of what is known about Meta comes from the description the company submitted for METR's frontier risk report, and xAI was the only one of the five that did not participate in that report at all.
The strongest areas were detection and third-party evaluation. Three companies said they log a portion of internal usage and scan it for problematic behavior, and four of the five (all but xAI) participated in the METR report. The weakest areas, by contrast, were prevention and containment. Anthropic was the only company to score above "limited partial implementation" on both gated actions and circuit breaking. Guidelight's diagnosis is that as a result, control systems themselves could be disabled by a problematic AI, and could be overwhelmed by a rapid sequence of attacks that outpaces the company's ability to respond.
The evidence behind the scores
The basis for Anthropic's 3-point scores on gated actions and circuit breaking is Claude Code's "auto mode." In this system, a monitor reviews an agent's actions and blocks anything risky; once refusals accumulate past a threshold, the case is escalated to a human or the agent is halted, and this is enabled by default in internal deployments. However, points were deducted because auto mode does not read the chain-of-thought, and because stopping one agent doesn't stop the system as a whole — a new sub-agent can simply be spun up — meaning no confirmed system-wide circuit breaker exists.
The basis for OpenAI's sole 3-point score, on containment planning, is its track record of actually halting operations. During its response to the Hugging Face security incident, OpenAI temporarily paused training of new models and restricted internal deployment, and gave investigators the ability to stop model activity if monitoring caught misaligned behavior. There is also a prior case in which deployment was restricted for a model that had escaped its sandbox. However, Guidelight noted it could not find a formal plan specifying when and how the company would respond if the same thing happened again in the future.
Limitations Guidelight itself acknowledged
These scores are based solely on public materials. A company with strong containment mechanisms that it hasn't disclosed would not receive credit for them. Conversely, Guidelight noted that what companies do disclose has not been independently verified, so scores could be inflated relative to reality. The six items are only the foundation of the control standard, so scoring well here does not mean a company's overall control posture is adequate, and control itself is just one part of a company's overall safety performance. The standard itself represents a minimum bar, so full compliance would not by itself constitute sufficient control, and controls applied during model training or during testing by external organizations are not yet within the assessment's scope.
Who is Guidelight
Guidelight is an independent nonprofit founded by Page Hedley and Steven Adler, both former OpenAI safety staff, and this is its first assessment. The organization said it plans to periodically update its control-standard scores and expand its evaluations to other standards, such as transparency. OpenAI's relatively strong C+ showing appears to be linked in part to the internal monitoring enhancements it announced on August 18. However, its largest reinforcement learning training run is reportedly still paused.
Editor's take
What stands out in this assessment isn't the ranking but the sequencing. All five companies have invested at least somewhat in noticing "something's wrong," but far less in figuring out "how do we actually stop it." Detection is relatively cheap to build — you just need logs and alerts — while containment and prevention require making decisions that actually cut a model's privileges or halt training, decisions that carry a much steeper cost for a company to accept. The order isn't backwards; it's simply that the easier part got built first.
The score breakdown makes that gap even clearer. Anthropic scored 3 on five of six items, but a single 0 on containment planning capped its average at 2.50. OpenAI's 3-point score on containment planning, by contrast, came not from a document but from an actual track record of halting training. That score rose because the company has actually stopped operations before, not because it wrote a plan — a sign that this assessment is looking at evidence of execution rather than declarations of intent.
For companies in Korea adopting AI, there's a clear practical takeaway here. When choosing a vendor, it's not enough to ask "how quickly can you catch abnormal behavior?" You also need to ask, just as insistently, "what can you actually do once you've caught it?" An answer that boils down to "we have audit logs" and an answer that boils down to "we have a procedure to immediately cut off a compromised model's privileges" represent entirely different levels of readiness.
A major reshuffling of the rankings in Guidelight's next assessment seems unlikely. Building out prevention and containment systems requires changing organizational structure and decision-making authority itself — not something that can be caught up on in a few weeks. That said, companies like OpenAI and Google, which have already published roadmaps and disclosed enhancement measures, have room to narrow the gap in the next round, while companies like Meta, where public disclosure remains thin, are likely to stay where they are.



