
이미지: The Decoder
Summary
- Anthropic's classifier designed to block biological and chemical weapons risk information was inactive for roughly 11 months, from May 2025 to April 2026
- During that period, about 50,000 external human-feedback contractors ran roughly 133 million model interactions without the filter in place
- Anthropic said its internal investigation found no confirmed cases of actual misuse, and it has since tightened its contractor vetting requirements
- 필터 비활성 기간
- 2025년 5월~2026년 4월 (약 11개월)
- 노출된 계약자 수
- 약 5만명
- 필터 없이 실행된 상호작용
- 약 1억3300만건
- 내부 조사 결과
- 실제 오용 사례 미확인
- 사후 조치
- 외부 계약자 심사 요건 강화
- 관련 사안
- Fable 5 분류기는 연구자 항의로 최근 완화
- 출처
- Anthropic Redacted Risk Report
133 million interactions passed through unfiltered
According to a safety report published by Anthropic, an internal classifier system designed to filter out risky information related to biological and chemical weapons was inactive for about 11 months, from May 2025 to April 2026. During that window, roughly 50,000 external contractors who provide human feedback interacted with the model without the filter active, amounting to approximately 133 million conversations — a volume that would take decades for a single person to work through, even at eight hours a day without a break.
What was turned off
The classifier is a safeguard meant to prevent the model from being exploited to extract dangerous knowledge related to producing chemical or biological weapons. According to the report, this contractor pool was vetted not by Anthropic itself but by external vendors using their own criteria, and in many cases that vetting process appears to have been inadequate. The fact that the filter remained off for nearly a year points to a blind spot in the internal system meant to verify that safeguards are actually functioning.
Anthropic said its own investigation found no evidence of actual misuse during this period. However, this is a conclusion reached only through a retrospective review — it does not change the fact that more than 100 million unfiltered conversations accumulated in real time while the filter was disabled.
Why this matters
Anthropic CEO Dario Amodei has reportedly identified AI-enabled chemical and biological weapons development as a greater threat than cyberattacks. That the company itself left the filter meant to catch such risks disabled for nearly a year highlights a gap between stated safety policy and actual operations. This case underscores an argument that's likely to gain traction: for companies working with frontier models, classifiers like this one can't simply be deployed and forgotten — they require a separate, ongoing monitoring system to confirm they remain active and functioning correctly.
Anthropic's response
Anthropic said that after identifying the issue, it tightened its vetting requirements for external contractors. The report does not specify whether the company changed the practice of delegating vetting to vendors altogether, or simply raised the bar for existing criteria. Anthropic disclosed the matter as part of its Responsible Scaling Policy, through the periodic risk reports it publishes.
Some filters were also loosened
In the same report, Anthropic said it had recently relaxed a classifier applied to its "Fable 5" model, after researchers complained that the filter was operating too aggressively and blocking legitimate research. The report thus contains, side by side, a case where a filter sat off for a year without anyone noticing, and a case where a filter was working so aggressively that it drew complaints.
Editor's view
The point worth dwelling on here isn't the conclusion that "no misuse occurred" — it's the fact that no one noticed for 11 months. A safety filter's real test isn't the moment it's deployed, but everything that happens after. Frontier model companies compete to announce new classifiers and new guardrails, but the operational systems that verify those safeguards are actually switched on have received comparatively little attention. This incident looks like a case study in exactly how that gap can show up.
Watching multiple companies announce safety improvements around the same time, a familiar pattern keeps repeating. When a new model launches, safety cards and classifier performance figures are unveiled with fanfare — but reports confirming that those same classifiers are still functioning months later are far rarer. The fact that an entire contractor pool was exposed without a filter, and that this only surfaced through a retrospective report, reveals a structural reliance on after-the-fact audits rather than real-time monitoring.
For companies in Korea deploying AI models or running their own fine-tuning, the lesson here is clear: attaching a safety filter or content classifier isn't the end of the job. A separate checklist is needed to periodically verify that the filter is actually still active. When vetting is outsourced to external vendors, it's safer to run independent sample checks rather than simply trusting the vendor's own standards.
In the coming weeks, other frontier model companies are likely to face pressure to disclose similar internal review results. Safety reports may shift away from being mere performance scorecards and start including details on what was disabled, and for how long.



