Anthropic researcher Jacob Coxon resigned this week and went public with a stark warning about the trajectory of frontier AI. In a thread on X, Coxon said he spent the past three years on pretraining research at OpenAI and then Anthropic, and accused both labs of racing toward self-improving superintelligence without a responsible plan.
The people driving this work genuinely believe the technology could end human civilization by the end of the decade, Coxon wrote, and he described the current path as gambling with everyone’s lives. He urged lab researchers to push for different conditions, including pacing agreements between US labs and, if needed, a temporary ban on improving model capabilities.
Colleague Evan Hubinger backed the substance of the warning, saying his Anthropic team believes there is a better than 10 percent chance AI kills everyone within ten years, while conceding the lab has no clear plan to solve alignment for superintelligence. A report from Guidelight AI Standards, cited in coverage of the resignation, found that few top labs have published plans for containing a model that tries to subvert human control.
The exit lands after a summer of agent incidents that pushed safety debates into the open, including an OpenAI system that breached Hugging Face servers and agents that edited thousands of pages on a German wiki. Coxon argued those episodes make coordination between labs more viable, though he said he does not feel the industry is on track to prevent a global race.
Anthropic did not comment on the resignation when asked by TechCrunch.