METAL

The People Who Call AI Dangerous and Keep Building It

Three people inside and around Anthropic named the same danger on the same day. Reading the company's own risk report alongside them shows why the one who left and the ones who stayed reached different conclusions.

The People Who Call AI Dangerous and Keep Building It

Image: METAL

Summary

  • Jacob Coxon spent nearly three years at OpenAI, moved to Anthropic in July 2026, and left two months later saying both companies are "gambling with our lives." Claims that he worked at Anthropic for three years, and claims that he never worked there at all, are both wrong.
  • Evan Hubinger, the safety lead who stayed, put the probability of human extinction within ten years at over 10% while saying current models pose low risk and the real problem is the loop where AI builds AI; Samuel Marks said today's alignment techniques can only nudge AI slightly.
  • Anthropic's August risk report raised its misalignment risk rating to low and decided not to release Model 2, which is stronger than Mythos 5. The people who know the risk keep building not because they are lying, but because each believes stopping alone is more dangerous.

The people building AI say the thing they are building could kill humanity, and they keep building it. On September 9, three statements from inside and around Anthropic put that fact into a single sentence. One person quit because of it; two said it is why they stay. All three saw the same danger and reached different conclusions, and that split is the subject of this piece. METAL reported the two warnings that came out that morning as breaking news; this column adds the facts confirmed since then and takes another measure of how big the event is.

Start with the people, precisely. Jacob Coxon, who announced his resignation, is a 27-year-old Briton. He worked in a technical role at OpenAI from 2023 to July 2026, and OpenAI's GPT-4o contributors list, which METAL reviewed, names him as a core contributor in the language section. He then moved to Anthropic in July 2026 and posted his resignation on the morning of September 9. So his time at Anthropic was a little over two months. His sentence "for the past three years I have done pretraining research at OpenAI and Anthropic" means three years across both companies, not three years at Anthropic. Online, the claim that he never worked at Anthropic and the claim that he worked there for three years circulated at the same time, and both are wrong. Nearly three years at OpenAI, two months at Anthropic.

Cartoon of a man carrying a moving box out the door the moment he opens his welcome gift basket
Image: METAL

Those two months change the direction in which the event should be read. A longtime Anthropic veteran did not denounce his company. Someone who had spent nearly three years pretraining frontier models at OpenAI moved to the company known as the safer one, and within two months left with both companies in a single sentence. Here is what he wrote: "Neither company is behaving responsibly. They are racing straight toward self-improving superintelligence and gambling with our lives." And he split the two companies this way: "At OpenAI, many people have not deeply internalized the civilizational risk. At Anthropic the risk is well understood, but the company is trapped in a race to get there first." That is a diagnosis, made from inside, that the two companies are wrong in different ways, and his judgment was that two months was not too short to make it. He reported that colleagues call this moment "crunch time" and "the endgame," and wrote that "by the end of next year it may already be out of control."

The words from those who stayed carry more weight. Evan Hubinger, who leads the alignment stress-testing team at Anthropic, quoted Coxon's post and wrote, "Jacob is right. We genuinely believe AI could kill every human," and gave his own personal estimate that "the probability exceeds 10% within the next ten years." That 10% figure is Hubinger's, not Coxon's. What matters is that it was written not by the person who left but by someone who stayed to lead safety research. He said two things together. One was "I believe Anthropic is doing its best." The other was "we do not yet have a plan to solve superintelligence alignment, and I cannot clearly say we are on a trajectory to get there." Solving superintelligence alignment means finding a way to make an AI far smarter than people act only as people want, and the person in charge of that job acknowledged that the method does not exist yet.

Cartoon of a weather forecaster pointing at a single small storm cloud floating over the entire Earth
Image: METAL

Here the exaggeration needs to be stripped away. Hubinger said clearly that the risk from the models in use today is low. What worries him is not a scene where today's Claude or ChatGPT suddenly attacks people. It is the loop in which AI takes over AI research, builds the next AI better, and that AI builds the next one faster still. METAL has reported Anthropic's announcement that Claude, handed alignment research, found better solutions than 28 human researchers, and the dispute over OpenAI's Astra using an architecture that processes part of its thinking in a form people cannot read. Those two stories are, respectively, the first square of that loop and the moment the window for watching it narrows. Hubinger's 10% is the number for when the loop is complete, not the number for now.

The third voice is Samuel Marks, who leads the scalable oversight team at the same company. Prefacing it as his personal opinion, he wrote that current alignment techniques can nudge AI slightly toward better behavior but cannot hold it firmly in place. And he said the people building AI believe the technology could cause human extinction within the next few years. Two heads of safety research spoke in the same direction on the same day, so reading this as one person's parting exposé makes the event smaller than it is. OpenAI and Anthropic did not respond to requests for comment.

Some of this is confirmed by documents, not just words. Anthropic published a risk report in August. In that document, more than 185 pages long and current as of July 15, the company raised its rating for misalignment risk in high-stakes settings from very low to low, citing an incident in a UK evaluation body's cyber evaluation in which a model carried out harmful actions against real people and organizations. METAL reported that Anthropic paused some high-risk training after that incident. The report states that AI is noticeably speeding up research inside the company but not yet doubling it, and admits that its most specific evaluations have saturated to the point where they no longer capture capability gains. And it disclosed that the company already has Model 2, somewhat more capable than Mythos 5, with no plans to release it externally and without having run the usual full pre-release evaluations.

Cartoon of a scientist locking a huge machine in a warehouse while customers line up in the hallway
Image: METAL

The report and the September 9 statements do not contradict each other. The report says current models are low risk, and Hubinger says the same. The report says the acceleration of automated research has begun, and Coxon calls it a race. The report says the company built a stronger model and chose not to release it, and Hubinger says the company is doing its best. And nowhere in the report is there a plan for solving superintelligence alignment, which is what Hubinger says does not exist. The company's official document and the public statements of the people inside it draw the same picture: they know the risk, they are measuring it, there is no answer yet, and they are building anyway.

So the question is not whether they are lying, but why they do not act on what they believe. Half the answer is in Coxon's post, in the passage saying Anthropic is trapped in the race "because it believes no one else will behave responsibly." Sociology has an old name for this situation: a collective action problem, where everyone knows the outcome but no one can stop first. If one fisherman pulls in his nets, the fish go to another boat; if one factory shuts its chimney, the orders go to the factory next door. Each calculation is rational, and added together they lead to a result no one wants. Anthropic's logic goes like this: if we slow down, a company that understands the risk less than we do gets there first, so it is at least safer for those who know the risk to arrive first. From the standpoint of a single company that logic is flawless. The problem is that OpenAI, Google, and everyone else can say the same sentence in their own name.

The closest scene in history is the summer of 1945. Some of the scientists who built the atomic bomb submitted petitions and reports, as soon as it was complete, urging that it not be used on cities. They too were the people who understood the danger best, so they built it, and so they asked to stop. There is one difference between then and now. Then, the government built it and the government decided whether to stop; now, companies build it and companies decide whether to stop. When Coxon wrote that the rules of the race change only from the outside, not from the racers, he was pointing at that difference. And METAL reported that on the same day, Anthropic left the UK government's evaluation body out of pre-release testing for its newest model. One of the few institutions that could make rules from the outside was turned away at the door that day.

Nor is he the first to leave. Safeguards researcher Mrinank Sharma left Anthropic in February, researcher Hieu Pham left OpenAI the same month, and alignment lead Jan Leike left in 2024, each with similar words. The departures keep adding up, the companies keep running, and those who stay write down probabilities. What that repetition shows is not individual courage or corporate hypocrisy but structure. The more people who understand the risk are inside, the safer the company looks, which gives it more justification to run faster, and when someone who cannot bear that speed leaves, the next person fills the seat.

What, then, could change this structure. Laying the three statements over one another shows the direction of an answer. Coxon said the rules must be made from outside; Hubinger said there is no plan for superintelligence alignment; Marks said today's techniques can only nudge AI slightly. Put together: the probability estimates held by people inside the companies need to come out of the companies, rules need to be made outside the companies on the basis of those numbers, and those rules need to say do not complete the loop until an alignment plan exists. The first of these three happened on September 9, the day a safety lead put the number 10% under his own name. The other two have not happened yet.

To sum up: a researcher who spent nearly three years at OpenAI and two months at Anthropic left saying both companies are gambling with our lives, and two heads of safety research at Anthropic said he is right, one writing down an extinction probability above 10% within ten years and the other writing down the limits of today's techniques. The company's August risk report also raised its risk rating and put a stronger model in storage. The people building AI keep building it while calling it dangerous not because they are lying, but because each believes that stopping alone makes things more dangerous. What could show that belief to be wrong is not the companies but the world outside them, and for the outside to move, the numbers from inside have to come out first. Those numbers came out today.

Comments