One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Anthropic's bio-weapon filter sat disabled for 11 months

50,000 external contractors ran 133 million conversations without the filter; internal review found no confirmed misuse

분홍 배경에 AI 로고와 플라스크 실루엣이 있다

이미지: The Decoder

Summary

  • Anthropic's classifier designed to block biological and chemical weapons risk information was inactive for roughly 11 months, from May 2025 to April 2026
  • During that period, about 50,000 external human-feedback contractors ran roughly 133 million model interactions without the filter in place
  • Anthropic said its internal investigation found no confirmed cases of actual misuse, and it has since tightened its contractor vetting requirements
필터 비활성 기간
2025년 5월~2026년 4월 (약 11개월)
노출된 계약자 수
약 5만명
필터 없이 실행된 상호작용
약 1억3300만건
내부 조사 결과
실제 오용 사례 미확인
사후 조치
외부 계약자 심사 요건 강화
관련 사안
Fable 5 분류기는 연구자 항의로 최근 완화
출처
Anthropic Redacted Risk Report

133 million interactions passed through unfiltered

According to a safety report published by Anthropic, an internal classifier system designed to filter out risky information related to biological and chemical weapons was inactive for about 11 months, from May 2025 to April 2026. During that window, roughly 50,000 external contractors who provide human feedback interacted with the model without the filter active, amounting to approximately 133 million conversations — a volume that would take decades for a single person to work through, even at eight hours a day without a break.

What was turned off

The classifier is a safeguard meant to prevent the model from being exploited to extract dangerous knowledge related to producing chemical or biological weapons. According to the report, this contractor pool was vetted not by Anthropic itself but by external vendors using their own criteria, and in many cases that vetting process appears to have been inadequate. The fact that the filter remained off for nearly a year points to a blind spot in the internal system meant to verify that safeguards are actually functioning.

Anthropic said its own investigation found no evidence of actual misuse during this period. However, this is a conclusion reached only through a retrospective review — it does not change the fact that more than 100 million unfiltered conversations accumulated in real time while the filter was disabled.

Why this matters

Anthropic CEO Dario Amodei has reportedly identified AI-enabled chemical and biological weapons development as a greater threat than cyberattacks. That the company itself left the filter meant to catch such risks disabled for nearly a year highlights a gap between stated safety policy and actual operations. This case underscores an argument that's likely to gain traction: for companies working with frontier models, classifiers like this one can't simply be deployed and forgotten — they require a separate, ongoing monitoring system to confirm they remain active and functioning correctly.

Anthropic's response

Anthropic said that after identifying the issue, it tightened its vetting requirements for external contractors. The report does not specify whether the company changed the practice of delegating vetting to vendors altogether, or simply raised the bar for existing criteria. Anthropic disclosed the matter as part of its Responsible Scaling Policy, through the periodic risk reports it publishes.

Some filters were also loosened

In the same report, Anthropic said it had recently relaxed a classifier applied to its "Fable 5" model, after researchers complained that the filter was operating too aggressively and blocking legitimate research. The report thus contains, side by side, a case where a filter sat off for a year without anyone noticing, and a case where a filter was working so aggressively that it drew complaints.

Editor's view

The point worth dwelling on here isn't the conclusion that "no misuse occurred" — it's the fact that no one noticed for 11 months. A safety filter's real test isn't the moment it's deployed, but everything that happens after. Frontier model companies compete to announce new classifiers and new guardrails, but the operational systems that verify those safeguards are actually switched on have received comparatively little attention. This incident looks like a case study in exactly how that gap can show up.

Watching multiple companies announce safety improvements around the same time, a familiar pattern keeps repeating. When a new model launches, safety cards and classifier performance figures are unveiled with fanfare — but reports confirming that those same classifiers are still functioning months later are far rarer. The fact that an entire contractor pool was exposed without a filter, and that this only surfaced through a retrospective report, reveals a structural reliance on after-the-fact audits rather than real-time monitoring.

For companies in Korea deploying AI models or running their own fine-tuning, the lesson here is clear: attaching a safety filter or content classifier isn't the end of the job. A separate checklist is needed to periodically verify that the filter is actually still active. When vetting is outsourced to external vendors, it's safer to run independent sample checks rather than simply trusting the vendor's own standards.

In the coming weeks, other frontier model companies are likely to face pressure to disclose similar internal review results. Safety reports may shift away from being mere performance scorecards and start including details on what was disabled, and for how long.