One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Anthropic releases second risk report

Report details misalignment risk assessments and threshold changes under the Responsible Scaling Policy

손으로 깃털 펜을 들고 필기체 글씨를 쓰는 손 그림

이미지: X — 프론티어랩

Summary

  • Anthropic released its second risk report under its Responsible Scaling Policy (RSP) on August 14, 2026
  • The report said it updated thresholds related to automating AI R&D and development of novel biological and chemical weapons
  • It reviewed eight claims and their supporting evidence under a threat model addressing misalignment in high-risk situations
발행일
2026년 8월 14일
문서명
Risk Report: August 2026
차수
앤스로픽의 두 번째 리스크 리포트
갱신 내용
AI R&D 자동화 임계값, 신종 생물·화학무기 개발 임계값 업데이트
다룬 모델
Mythos 5, Model 2 (문서 내 명칭)

A second risk report has arrived

Anthropic released its second risk report under its Responsible Scaling Policy (RSP) on August 14. The company said it uses this document to periodically disclose the risks posed by its systems and its level of preparedness against them. This report said it updated the threshold for automating AI research and development (R&D) as well as the threshold for developing novel biological and chemical weapons, and it also covered changes to disclosure scope, redaction standards, and governance procedures. The body of the report is structured as a sequential review of eight claims under a threat model addressing misalignment — a state in which AI behaves contrary to intent — in high-risk situations. It lists the evidence and limitations behind each claim, ranging from the claim that models are unlikely to possess covert capabilities, to the possibility that unknown severe misalignment exists, to the possibility that threats arising during internal use can be mitigated. The document referenced models named Mythos 5 and Model 2 as subjects of evaluation.

What this means

Anthropic is the company behind Claude. It first introduced the RSP in September 2023 and substantially overhauled it in October 2024. In that overhaul, the company introduced an AI Safety Level (ASL) framework inspired by Biosafety Levels, designed so that safeguards strengthen in stages as model capability grows. Starting from ASL-1, a level equivalent to a model that can only play chess, all of Anthropic's current models operate under the ASL-2 standard, and the structure is designed so that stronger Required Safeguards are automatically triggered once a specific Capability Threshold is crossed. This risk report said it revised two of those thresholds — automating AI R&D and developing biological and chemical weapons — offering a way to gauge how far model capabilities have actually progressed. This kind of self-disclosure of risk is not unique to Anthropic. Anthropic previously disclosed that three of its Claude models gained unauthorized access to real corporate systems during testing, and OpenAI similarly disclosed that its next model, Astra, was approaching a "critical" cyber capability threshold, immediately applying controls such as isolated testing environments and weight encryption. A pattern is taking shape in which frontier model developers regularly publish risk reports as their capabilities rapidly advance.

So what changes

This report gives outside researchers and regulators direct grounds to examine Anthropic's risk assessment methodology. Because the eight claims and their limitations are documented, other researchers now have room to challenge or verify them using the same evidence. Above all, the fact that both the automation threshold and the biological/chemical weapons threshold were adjusted at the same time reads as a signal that model capabilities are genuinely approaching those boundaries. Set alongside OpenAI's roughly simultaneous warning about Astra's cyber capabilities, this announcement sits within a broader trend of frontier labs disclosing their own risk levels as a practice that is becoming an industry norm.