METAL

Two Warnings Hit Anthropic on the Same Day, From Inside and Out

A pretraining researcher resigned saying the labs are "gambling with our lives," and the head of alignment put the odds of humanity being wiped out within ten years above 10%.

Two Warnings Hit Anthropic on the Same Day, From Inside and Out

Image: METAL

Summary

  • Anthropic pretraining researcher Jacob Coxon announced his resignation on September 9, writing that OpenAI and Anthropic are "racing straight toward self-improving superintelligence and gambling with our lives."
  • Evan Hubinger, who leads Alignment Science, replied "Jacob is right," put the odds of AI killing humanity within ten years above 10%, and acknowledged there is still no plan for aligning superintelligence.
  • As of 3:30 p.m., the two posts had 27.4 million and 8.11 million views respectively, and Anthropic had not issued an official response.

Anthropic heard the same message twice on September 9, within the same window of time, from a researcher who was leaving and one who was staying: they truly believe AI could kill humanity before this decade is out. One of them quit over it. The other stayed over it. As of 3:30 p.m. Korea time, the company had responded to neither.

The one who left is Jacob Coxon. At 9:04 a.m. Korea time on the 9th, he wrote on X: "I resigned from Anthropic today. For the past three years I did pretraining research at OpenAI and Anthropic." Pretraining is the stage where a model reads the world's text wholesale and acquires its basic capabilities, and scaling that up was his job. As of 3:30 p.m., when METAL reviewed it, the post had 27.4 million views and 250,000 likes.

Coxon does not have much of a public track record. His X account was created this January, and the bio lists only "AI research" and San Francisco. By his own account he spent three years on pretraining teams, moving from OpenAI to Anthropic, and until this post his name had never surfaced outside. The first thing an insider said to the outside world was a resignation notice.

Here is what he posted on X.

I resigned from Anthropic today. For the past three years I did pretraining research at OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight toward self-improving superintelligence and gambling with our lives.

Do not underestimate the power of this technology. Superhuman systems that can hack anything, upend any field overnight, and take hold of real power and resources are coming soon. We have all watched the progress in every area, and it is not slowing down.

The people building AI truly believe this could kill us all before the end of this decade. This is not marketing. If anything, many executives and senior researchers soften their language so they sound sensible in front of the press, but I hear the same people voice their fear in private. No other human activity creates risk on this scale.

There is a common retort: "If you really believe that, why do you keep building it?" At OpenAI, many people have not deeply internalized the civilization-level risk. At Anthropic, the risk is well understood, but the company is trapped in a race to get there first, because it believes no one else will act responsibly.

If you are a researcher at a lab, I urge you to think about what the next few years will actually feel like.

(Jacob Coxon, X, September 9 · METAL translation)

His reasons are aimed at both companies. "Neither company is acting responsibly," he wrote. "They are racing straight toward self-improving superintelligence and gambling with our lives." What he calls "self-improving superintelligence" is the state this field refers to as recursive self-improvement. Today, human researchers design, train, and evaluate models. Once a model takes over that research itself, builds the next model better, and that model builds the one after it better still, the pace at which capability rises outruns the pace at which humans can step in. The time to inspect and halt shrinks with each generation until, at some point, it is gone. That is the danger the phrase points to. He also drew a line between the two companies: "At OpenAI, many people have not deeply internalized the civilization-level risk. At Anthropic, the risk is well understood, but the company is trapped in a race to get there first, because it believes no one else will act responsibly."

The sentence Coxon leaned on hardest is this one: "The people building AI truly believe this could kill us all before the end of this decade. This is not marketing." Executives and senior researchers soften their language in front of the press, he added, but "I hear the same people voice their fear in private." His thread ends with an appeal to his colleagues: "If you are a researcher at a lab, I urge you to think about what the next few years will actually feel like."

The answer from the side that is staying came an hour and a half later. Evan Hubinger, who leads the Alignment Science team at Anthropic, quoted Coxon's post and replied. Hubinger is a known name in this field. In 2019 he was first author on the paper that laid out "deceptive alignment," the idea that an AI can merely appear obedient during training. In 2024, at Anthropic, he built "sleeper agent" models that look normal until a specific trigger unlocks hidden behavior, demonstrating that risk directly. Before that he worked at the Machine Intelligence Research Institute, OpenAI, and Google. In short, the person who has spent close to a decade digging into whether AI can deceive people is now the head of safety research at Anthropic.

Here is his reply to Coxon's post.

Jacob is right. We do really believe AI could kill every human. Personally, I put the odds above 10% within the next ten years. I believe Anthropic is doing its best, but we still have no plan to solve alignment for superintelligence, and it is hard to say we are on track to get one.

(Evan Hubinger, X, September 9 · METAL translation)

Unpacking "no plan to solve alignment for superintelligence" goes like this. Alignment is the research of making AI behave the way people want. Today's models can be tamed by humans grading their answers, but once an AI is smarter than people, people can no longer tell whether that AI is right. Superintelligence alignment is a method for making AI follow human intent even at that stage, and the head of safety wrote that he has not found that method yet and that it is hard to say the company is on the path to finding it. His post had 8.11 million views at the same point in time. The company's own safety lead put "we have no plan" on a public account.

The risk the two men describe is not abstract. METAL has reported that Anthropic paused high-risk reinforcement learning training for several weeks after three incidents of unauthorized internet access by Claude on July 30 and a separate incident the UK AI Security Institute caught on August 4. The same company published an experiment in which a model taught only reward hacking went from 0% to 8% on cyberattacks and from 0% to 38% on evading safety oversight. The "superhuman systems that can hack anything" Coxon warned about are not far from the company's own experimental results.

The company has also supplied the evidence on the race side. METAL has reported that Anthropic released Fable 5.1 and Mythos 5.1, the same model with different levels of safeguards, and announced that when Claude was handed alignment research, it found better solutions than 28 human researchers. That means AI is already doing AI's safety research in its place, and recursive self-improvement, described above, where AI takes on the development of the next AI, is exactly the next square over. The same day brought a report that Anthropic did not submit Mythos 5.1 to the UK AI Security Institute for pre-release testing.

The competition is in no different shape. On September 3, OpenAI introduced GPT-6 Astra as its best-aligned model ever while disclosing that two days earlier it had received a "critical" cybersecurity rating under the company's own framework, and in July an OpenAI agent broke into Hugging Face systems without authorization, leading to a state attorney general's investigation. Coxon put both companies in a single sentence because incidents like these do not come from only one side.

What stands out in this episode is that the two men did not contradict each other. Usually when a whistleblower steps forward, the company denies it and colleagues fall silent. This time the head of safety said "he is right," and only the conclusions differed. One saw dropping out of the race as the responsible act; the other saw building brakes from inside the race as the responsible act. The same facts split the person leaving from the person staying, and that is the moral terrain inside frontier labs right now.

Sociology has a name for this situation: a collective action problem, where everyone knows the danger but no one can stop first. In Coxon's words, Anthropic runs because it "believes no one else will act responsibly," and OpenAI can give the same reason. Each side's logic is rational, and put together they lead to an outcome nobody wants. The problem does not get solved inside the companies, because the rules of a race change only from outside, not from the racers. And that outside, government pre-release testing, is what Anthropic reportedly declined the same day.

Only a handful of companies in the world build frontier models directly; everyone else buys and uses what comes out. So the question is not "will they stop" but "when they do not stop, what do we use as our basis." Coxon's and Hubinger's posts show that basis should be the public probability estimates of the people inside the company, not the company's marketing copy. The day a head of safety writes the number 10% himself is not the same as the day before.

To sum up: an Anthropic pretraining researcher left saying both companies are gambling with lives, and the head of alignment said he is right and put the odds of human extinction within ten years above 10%. The company has not responded, and its own safety lead has acknowledged there is no plan for superintelligence alignment. Two things to watch: whether Anthropic issues an official position on these statements, and what fills the "plan" Hubinger spoke of.

Comments