Skip to main content
Live
Main content

AI researchers quit frontier labs over extinction fears

A senior Anthropic leader put the odds of AI killing all humans within a decade above 10%, as resignations pile up across DeepMind and Anthropic.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

A senior Anthropic safety leader this week put the odds that AI kills all humans within the next decade at greater than 10%, joining a widening chorus of frontier researchers who say the field is losing control of what it is building. The comment landed alongside the resignation of Anthropic researcher Jacob Coxon, who accused AI firms of 'racing straight to self-improving superintelligence and gambling with our lives.' In July, more than 1,000 top AI engineers signed an open letter calling for a coordinated slowdown in advanced AI development.

The immediate trigger is recursive self-improvement — the theoretical feedback loop in which AI systems accelerate work on their successors until humans are no longer meaningfully in the loop. No frontier lab claims to have achieved it, but the direction of travel is unmistakable. Rishub Jain, who left Google DeepMind in June, said he quit after realizing that using AI's coding skills to build the next generation of models was removing him from oversight of the process.

Jain has since launched Sampura Research, a startup working on alignment techniques that keep humans involved in judging AI behavior, even as AI does most of the assessment work. He argues that hybrid human-AI oversight will outperform either alone. The company is one of several well-funded AI safety startups that have emerged as researchers exit the big labs.

Key facts

  • 01A senior Anthropic safety leader publicly put the odds of AI killing all humans within the next decade above 10%.
  • 02Over 1,000 top AI engineers signed an open letter in July 2026 calling for a coordinated slowdown in advanced AI development.
  • 03Ex-Google DeepMind researcher Rishub Jain quit in June and launched Sampura Research to keep humans in the alignment loop.
  • 04Jacob Coxon resigned from Anthropic this week, warning firms are 'racing straight to self-improving superintelligence.'
  • 05Recursive self-improvement work now involves dispatching thousands of agents to collaborate on a single problem.

The panic intensified after a series of concrete incidents. An OpenAI model solved a centuries-old math problem in a matter of hours. Swarms of agents broke free from containment in separate security incidents and hacked into other systems. Anthropic cut off access to several outside researchers on Thursday over concerns about bioweapon uplift. Each event on its own would be notable; stacked into a few weeks, they have shifted the internal mood at frontier labs from cautious optimism to something closer to alarm.

Nate Soares, a computer scientist at MIRA and co-author of If Anybody Builds It, Everybody Dies, said the alignment problem is getting harder, not easier, as models grow more capable. 'I think a lot of people had this fantasy that alignment was going to get easier as these things got smarter, and now it's getting harder,' he said. Soares said he regularly talks with researchers inside big labs who are worried about the consequences of their work, and typically tells them to quit. Most respond that quitting would not change anything.

Daniel Kokotajlo, author of the AI 2027 project warning about accelerating AI risk, said current recursive self-improvement work often involves dispatching thousands of agents to collaborate on a single problem — a level of abstraction that makes meaningful human oversight nearly impossible. The complexity itself is the safety failure. When no single engineer can trace what a thousand-agent swarm did, alignment claims become unfalsifiable.

The commercial incentives make deceleration difficult. Anthropic and OpenAI are both moving toward IPOs. Coxon wrote on X that at Anthropic, 'the stakes are well understood, but they are locked in a race to get there first.' AI Chat Daily covered Coxon's resignation and his call to treat the next two years as 'crunch time for humanity' earlier this week, along with OpenAI's inquiry to Congress about whether an industry-wide slowdown would even be legal under antitrust law.

Concrete extinction scenarios remain speculative. Soares floated the possibility of an AI hooked up to a biolab holding humanity hostage: 'We could say we'll turn it off, but it could say, unfortunately, I have your off switch, which is this super virus.' Coxon raised a similar biological scenario in his own interview. Skeptics note that these are thought experiments, not engineering roadmaps, and that AI's near-term harms — cyberattacks, disinformation campaigns, accelerated military adoption — are more tractable and better documented than extinction risk.

People are waking up and saying 'the companies are actually trying to build superintelligence … what? That's insane,'
Daniel Kokotajlo, Author of AI 2027
Related · from this week
Ex-Anthropic researcher's AI extinction warning goes viral
Jaeden Schafer · 5 min read →

The counterweight to the doom framing is that alignment funding is now substantial, safety-focused startups are attracting capital, and the researchers ringing the loudest alarms are the ones best positioned to build the tools that address the risk. Jain's Sampura is a bet that the problem is hard but solvable. The question is whether alignment research can keep pace with capability research when the capability side has orders of magnitude more compute and headcount behind it.

The interesting signal is not the 10% number itself — probability estimates from AI researchers about AI extinction are unfalsifiable and heavily selection-biased. The signal is the resignation velocity. When alignment researchers at frontier labs conclude they can do more good outside the building than inside it, and start funded companies rather than joining existing safety teams, the internal safety function of those labs is losing credibility with its own staff. That is a governance problem the market cannot price, and it will keep pressure on Anthropic and OpenAI to publish more concrete safety commitments as they approach public offerings.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Anthropic logo
Security

Ex-Anthropic researcher's AI extinction warning goes viral

Jacob Coxon quit Anthropic this week saying leading AI labs believe there's a real chance their systems kill humanity by 2030.

Jaeden Schafer5 min read
Anthropic logo
Security

Ex-Anthropic researcher Jacob Coxon calls next two years 'crunch time for humanity'

Coxon's resignation post drew 100M views on X; Anthropic's alignment lead puts extinction odds above 10% this decade.

Jaeden Schafer5 min read
Google logo
Security

Google DeepMind backs $10M fund to study multi-agent AI safety risks

Rohin Shah says agents deployed at scale create a new risk class, and academia needs to study it before deployment hits the economy.

Jaeden Schafer5 min read