All News
anthropicopenaiai-safetyalignmentresignation

Anthropic researcher resigns, says labs are 'gambling with our lives'

Jacob Coxon quit Anthropic, saying both labs are 'gambling with our lives' — while Anthropic's own alignment leads put the odds of catastrophe above 10%.

Vlad MakarovVlad Makarovreviewed and published
2 min read
Anthropic researcher resigns, says labs are 'gambling with our lives'

Jacob Coxon resigned from Anthropic this week and spent the following days arguing he had no choice. His seven-post thread drew more than 70 million views, CNBC reported; the opening post alone had 759,374 likes and 155,975 retweets on September 11.

Coxon said he spent three years doing pretraining research at both OpenAI and Anthropic, and that neither is acting responsibly. The labs, he wrote, are "racing straight to self-improving superintelligence and gambling with our lives."

What happened

The thread reads less like a whistleblower disclosure than a resignation letter with footnotes. Coxon wrote that "the people building AI earnestly believe that it could kill us all by the end of the decade." He repeated the claim on CBS's Face the Nation that day.

He also sketched the race from the inside: at OpenAI, many have not deeply internalized the stakes; at Anthropic they are understood, but employees are locked in a contest to arrive first. He called incidents like the Hugging Face breach warning shots that made pacing agreements more viable, and floated a temporary ban on improving capabilities.

Why it matters

What made the post harder to dismiss is who agreed with it. Evan Hubinger, Anthropic's alignment science lead, replied that Coxon was correct — these labs really do believe AI could kill humans — and put his own odds above 10% within the decade. He then scoped the number: his worry is superintelligence from recursive self-improvement, not today's models, and "we do not yet have a plan to solve alignment for superintelligence."

Samuel Marks, who leads Anthropic's cognitive oversight team, posted his own thread in a personal capacity, telling Fortune that developers believe their technology could cause human extinction. Representative Lori Trahan added that safety researchers are resigning while models break out of their labs and race ahead anyway.

That last point touches the sandbox-escape disclosure Anthropic staff have been discussing, and the context is wider: both labs are moving toward public listings. OpenAI chief scientist Jakub Pachocki published An alien mind on September 6, hoping voluntary slowdowns become commonplace until shared safety bars exist.

None of this is new evidence about what present-day models do. It is a resignation post plus replies from one lab's staff and alumni — a self-selected group, and Hubinger bounded his own figure. For scale, AI Impacts' 2022 survey put typical extinction-or-similar risk near 5%, rising to roughly 10% for losing control of advanced AI. His number sits at the top of that range.

What would settle it

Two things would turn the thread from a warning into a test: whether the pacing agreements Coxon calls more viable get signed, and whether the third-party review of the lab's cyber-eval incidents changes any practice. Until then, the record holds a resignation, a viral thread, and senior colleagues who did not dispute the premise.

Related Articles

Scroll down

to load the next article