Anthropic researcher quits, warns AI could cause extinction by 2030
Transformative AICoxon announced his resignation from Anthropic on Tuesday, 8 September 2026, in a multi-part thread on X under the handle @hilbertspaess, and the post drew tens of millions of views within hours. He accused both companies of "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." According to the Washington Post, he resigned over concerns that the company and the broader artificial intelligence industry were "gambling with our lives" by racing to build AI with self-improving superintelligence.
Coxon, a University of Cambridge graduate, spent three years in pretraining research, first at OpenAI from 2023 until July 2026, where he was among the core contributors to GPT-4o, then at Anthropic. He warned that near-future systems would become "superhuman" systems capable of hacking "anything" and acquiring "real power and resources". He argued the danger was already understood inside the labs themselves: "At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk." He pointed to a reported security breach involving OpenAI agents and Hugging Face as what he cited as a "warning shot," and called for coordination among AI labs plus a temporary ban on improving model capabilities.
His warning did not stand alone. Evan Hubinger, Anthropic's Alignment Science Lead, publicly backed him, writing that Coxon was "correct" that some researchers genuinely believe advanced AI could pose an existential risk, and stating that he personally believes there is a greater than 10 percent chance that AI could cause human extinction within the next decade. Hubinger also said, according to TechCrunch, that the risk from current models is low, but that the fear compounds with "superintelligence arising from recursive self-improvement," which is "happening faster than we thought". Samuel Marks, who leads scalable oversight at Anthropic, reportedly echoed the concern, suggesting that seniority within the labs correlates with greater alarm rather than less.
The episode lands amid a wider push for AI safety coordination. Hubinger was also one of roughly 1,400 AI researchers who signed an open letter called "Pacing the Frontier" in July, and, per one tally, more than 1,000 AI researchers have now signed onto a statement calling for global coordination on a way to hit the brakes if things start slipping. In Washington, Sen. Bernie Sanders and Rep. Greg Casar reportedly introduced legislation to ban superintelligent AI outright, and want to freeze development until real safety rules exist. The debate over Coxon's forecast is not settled: as CoinDesk notes, skeptics dispute such extinction forecasts, but research suggests AI is already affecting the labor market, with entry-level employment down nearly 20% in U.S. sectors most exposed to the technology.
Anthropic has long marketed itself as the safety-conscious counterweight to rivals racing toward more powerful models, describing itself, per the San Francisco Chronicle, as a public benefit corporation dedicated to developing AI while mitigating its risks. That positioning sits uneasily alongside reports, cited by Yahoo News, that Anthropic is reportedly preparing for a huge IPO that could value it close to $2 trillion, while telling investors it takes safety more seriously than anyone else in the room. Coxon, for his part, said the company's most consequential model decisions play out informally, describing debates over its riskiest systems as happening on "Slack and engineers' laptops in San Francisco".