Industry·4 min read
By BitsMindsSource: TIME / TechCrunch

Anthropic Alignment Lead: Over 10% Chance AI Kills Us All

Evan Hubinger, who leads alignment stress-testing at Anthropic, says he personally puts the odds of AI killing all humans at over 10% within the next decade — and that the company "does not yet have a plan to solve alignment for superintelligence". He was backing Jacob Coxon, a pretraining researcher who quit on September 8 accusing Anthropic and OpenAI of "gambling with our lives".

Anthropic Alignment Lead: Over 10% Chance AI Kills Us All
Share:

Evan Hubinger, who leads alignment stress-testing at Anthropic, wrote publicly this week that he personally puts the odds of AI killing every human being at better than one in ten within the next decade. "Jacob is correct here — we really do earnestly believe AI could kill all humans," he posted. "I personally think it is >10% within the next decade." He added that Anthropic "is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

That is an unusual sentence to read from the person whose job is to pressure-test whether his employer's models are safe. Senior figures at frontier labs have signed open letters about extinction risk before, in language general enough to mean very little. A named probability, attached to a decade, from a sitting department lead who is not resigning and not hedging, is a different kind of statement.

It surfaced as a reply. On Tuesday, September 8, Jacob Coxon — a 27-year-old British researcher who spent roughly three years on pretraining at OpenAI before joining Anthropic earlier this year — posted a seven-part thread on X announcing his resignation. "Neither company is acting responsibly," he wrote. "They are racing straight to self-improving superintelligence and gambling with our lives." His central claim was not a prediction of his own but a description of the room: the people building this technology, he said, "earnestly believe it could kill us all by the end of the decade", and say so privately in terms far starker than they use in public.

Coxon's specific ask was narrow and technical. He argued that the leading labs should stop accelerating recursive self-improvement — the practice of pointing powerful internal models at the work of building the next generation of models. That is the loop he thinks removes the human brake, and BitsMinds has covered how much capital is now chasing exactly that capability. "It's obvious that things are speeding up," he told TIME, "and they're not under control."

Resignation-in-protest is a familiar genre in this industry, and it is easy to discount. Hubinger's reply is what makes the episode hard to file away, because it removes the obvious rebuttal. A departing employee can be waved off as disgruntled, or as someone who was never close enough to the safety work to judge it. The person running alignment stress-testing at the company Coxon accused is neither, and he did not push back — he confirmed the premise and supplied a number.

Two things are worth holding onto about that number. It is a subjective credence, not a measurement: there is no experiment behind ">10%", and Hubinger is describing a belief, not a finding. But it is also not a rhetorical flourish, because he pairs it with a concrete operational admission — no plan for superintelligence alignment, and no clear track toward one. Those are claims about the state of the work, and they are checkable in a way a doom estimate is not.

Reaction split along the lines it usually does. Technology journalist Taylor Lorenz dismissed the thread as "sanctimonious doomer posting", a common read that treats existential warnings from lab insiders as a marketing posture that makes the product sound more powerful than it is — a charge Coxon pre-empted by insisting the private conversations are worse than the public ones. Anthropic and OpenAI did not respond to press requests for comment on either the resignation or Hubinger's estimate.

What the exchange documents is a company that has spent the year committing hundreds of billions of dollars to compute while the person auditing its safety story says out loud that the hardest problem remains unsolved and that the timeline may not wait. Those two facts are not contradictory — Anthropic's stated bet has always been that it is safer for a safety-focused lab to be at the frontier than to cede it — but this is the first time the tension has been narrated from inside the alignment team, with a number attached.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles