Anthropic built its entire brand on being the "safety-first" alternative to OpenAI. Its mission statement promises AI that remains "helpful, honest, and harmless." Its research teams frame their work around ensuring advanced systems stay aligned with human values as they grow more capable.
On September 8, 2026, the lead of Anthropic's own alignment stress-testing team publicly endorsed a departing colleague's scorching resignation letter, and went further. Evan Hubinger stated he personally believes there's a greater than 10% chance AI could kill all humans within the next decade. More damning still, Hubinger admitted his own company, the one founded to solve this exact problem, "does not yet have a plan" for superintelligence alignment.
This isn't an external critic sounding alarms from the outside. It's the insider at the guard post saying the guard post has no walls.
The Resignation That Broke the Silence
The story begins with Jacob Coxon, who resigned from Anthropic after three years doing pretraining research across both OpenAI and Anthropic. In a public thread on X, Coxon accused both companies of "racing straight to self-improving superintelligence and gambling with our lives."
Coxon's trajectory matters. He's not an outsider looking in with academic concerns. He's someone who's seen the inner workings of the two most powerful AI labs on Earth and concluded that both are on an unsafe path. His thread, posted September 8, 2026, drew immediate attention across the tech world and beyond.
The significance of his resignation letter lay not in its novelty but in its source. People who work on pretraining at frontier labs are not typically inclined to burn bridges. They're well-compensated, deeply embedded in the culture, and surrounded by colleagues who share their technical interests. When someone in that position walks away and publicly condemns the entire enterprise, it carries weight.

The Insider's Admission: Hubinger's >10% Claim
Hubinger's response came quickly. As lead of Anthropic's alignment stress-testing team, he occupies a role specifically designed to find alignment failures before they become catastrophic. His public endorsement of Coxon's concerns was notable in itself. His personal probability estimate was the real headline.
"I believe there is a >10% chance AI could kill all humans within the next decade," Hubinger wrote, according to consistent reporting from BBC, CNBC, and IGN.
But the most striking admission came next. Hubinger acknowledged that Anthropic "is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
The attribution matters here. This isn't a general industry critique from a detached observer. Hubinger is speaking specifically about his own employer, the lab that positioned itself as the responsible alternative to OpenAI. If Anthropic lacks a coherent internal strategy for superintelligence alignment, the implications for the rest of the field are grim.
The Numbers in Context: How Big Is >10%?
A >10% chance of irreversible human extinction within ten years is a staggering risk by any standard. To put it in visceral terms, it's roughly comparable to the odds of rolling a 1 on a standard die, except the consequence is unrecoverable.
This figure aligns with existing expert consensus. A 2022 survey of AI researchers found a majority believed there is a 10% or greater chance that human inability to control AI causes an existential catastrophe. In 2023, hundreds of experts signed a statement on mitigating AI extinction risk. Not all researchers share this estimate, some place the risk far lower, but the fact that a sitting safety lead subscribes to the pessimistic end of the spectrum is itself notable.
Some analyses have suggested Hubinger's >10% figure is roughly 1.8 times larger than the biggest previously published extinction-risk estimate. That particular comparison comes from a low-authority source and should be treated cautiously. The real story isn't the precise arithmetic. It's that a leading lab's safety lead personally aligns with the field's worst-case assessments rather than dismissing them.
This is not a scenario about job displacement or economic disruption. Hubinger is talking about an unrecoverable outcome.

The Ironic Contradiction: 'Safety-First' Anthropic and the Race to Self-Improving AI
The core tension here is difficult to overstate. Anthropic positions itself as the responsible lab, the one that will get AI safety right where others might cut corners. Yet its own public research tells a more complicated story.
Anthropic's work on recursive self-improvement, documented in research like "When AI builds itself," describes a path toward AI systems that autonomously design their successors. In practice, this means AI systems that can modify their own code or architecture to become more capable without direct human intervention. If such systems eventually take over their own development and improve themselves faster than humans can keep up, the alignment problem compounds exponentially. You can't simply "fix it later" if the AI is advancing beyond human oversight capacity in the meantime.
Hubinger's admission lands in this context with uncomfortable precision. The lab is simultaneously researching the very capability that makes alignment most difficult while admitting it has no plan to solve that alignment.
Coxon's resignation letter accused both OpenAI and Anthropic of the same fundamental race. That framing suggests the "safety-first" branding may be more marketing than operational reality. Whatever the cultural differences between the labs, the trajectory looks similar from the inside.
What This Means, and the Credibility Problem
The story resonated far beyond AI policy circles. BBC, CNBC, IGN, and numerous other outlets covered it within 24 hours. The reason is straightforward: when the person responsible for stress-testing alignment at the industry's flagship safety lab says there's a double-digit chance of extinction and no plan exists, that's news.
A verification caveat is worth noting. The primary sources are X posts sitting behind login walls, so exact wording comes through secondary corroboration. However, the consistency across major outlets gives high confidence in the substance of what was said. The identity of the speaker is well-established: Hubinger's role as alignment stress-testing team lead is confirmed across multiple independent sources.
What does "no plan" actually mean in practice? It means no coordinated internal strategy for superintelligence alignment. No agreed-upon technical approach that the lab's leadership has committed to. No clear timeline for when alignment needs to be solved versus when superintelligence arrives. The research teams continue their work, but Hubinger's admission suggests that work hasn't coalesced into a unified plan capable of addressing the scale of the challenge.
If the safety lead at the safety-first lab is this candid about the gaps, the logical question becomes: what does the rest of the industry look like?
The Honesty Gap
Anthropic was founded to be the lab that got AI safety right. Its researchers are among the most thoughtful in the field, and its public commitment to responsible development has been consistent. And yet, the person responsible for stress-testing alignment says the odds of catastrophe are potentially double-digit percentages within a decade, and that his own company doesn't have a plan to prevent it.
The story isn't the probability number itself. It's that the people closest to the technology are the ones most willing to state the risk publicly. Coxon resigned and spoke out. Hubinger endorsed his concerns and added his own assessment. Both acted with apparent disregard for how their words might affect their careers or their companies' reputations.
Whether Hubinger's >10% estimate is too high or too low is almost beside the point. The question now is whether that honesty translates into action before the decade runs out.





Comments
Join the Conversation
Share your thoughts, ask questions, and connect with other community members.
No comments yet
Be the first to share your thoughts!