Anthropic researcher quits, warns AI race could ‘kill us all’; safety lead says extinction risk above 10%
text_fieldsA researcher at US-based artificial intelligence company Anthropic has resigned, accusing the firm and rival labs of racing toward self-improving superintelligence without adequate safeguards and “gambling with our lives”.
Jacob Coxon, who said he had spent the past three years doing pre-training research at both Anthropic and OpenAI, announced his departure in a series of posts on social media platform X on Wednesday. He claimed that many people building frontier AI systems “earnestly believe that it could kill us all by the end of the decade” and that no other human activity poses a comparable existential threat.
“These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources,” Coxon wrote. “We have all witnessed the progress in each of these domains, and progress is not slowing.”
Coxon contrasted his experiences at the two companies. At OpenAI, he said, many employees had not “deeply internalised the civilisational stakes”, which he argued helped explain why development continued despite known risks. At Anthropic, he claimed, the dangers were well understood but the company was “locked in a race to get there first”, operating on the assumption that no other player would act responsibly.
He described the decision to enter this “endgame” as a “hubristic gamble” that should not be made in private company chat channels, and said the industry was not on track to prevent a global race that might require drastic measures such as a temporary ban on advancing model capabilities. Coxon urged researchers to consider whether they were prepared to “kick off a superintelligent RL run without a rigorous understanding of its mind”, referring to large-scale reinforcement learning training that can push systems beyond human-level performance.
His posts drew more than 23 million views on X within hours. OpenAI lists Coxon as a core contributor to GPT 4o, a model released in May 2024. Anthropic, whose Claude series of assistants is used by enterprises and has been adopted by the US military, did not immediately comment on his resignation.
However, Evan Hubinger, Anthropic’s Alignment Science lead, responded publicly, saying Coxon was “correct” and that he and colleagues “really do earnestly believe AI could kill all humans”. Hubinger added that he personally assessed the probability at more than 10 per cent within the next decade.
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Hubinger wrote. He noted that while the company’s August report categorised risks from current models as low, he was “worried about superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought”.
Billionaire hedge fund manager Bill Ackman described Coxon’s allegations as “concerning”. The exchange has intensified debate over whether frontier AI labs are moving too quickly toward systems whose behaviour and long-term impacts they cannot reliably control.

