Skip to main content
Live
Main content

Anthropic safety lead puts AI extinction odds above 10% this decade

Alignment lead Evan Hubinger backed a departing colleague's warning that Anthropic and OpenAI are racing toward uncontrollable systems.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

Evan Hubinger, who leads an alignment science team at Anthropic, said on September 9, 2026 that there is a greater than 10% chance AI could kill all humans within the next decade. The statement came hours after Anthropic researcher Jacob Coxon resigned and accused the company and OpenAI of racing toward self-improving superintelligence without adequate safeguards. Both companies are approaching anticipated public listings while their own staff publicly question whether the technology can be controlled.

Coxon, who previously trained models at OpenAI before joining Anthropic, announced his departure in a post on X. He said neither lab is acting responsibly and that internal researchers hold serious extinction-level concerns while continuing to ship.

They are racing straight to self-improving superintelligence and gambling with our lives
Jacob Coxon, Former Anthropic researcher

Hubinger responded directly on X, endorsing Coxon's characterization and adding his own numeric estimate. He said Anthropic is trying its best but does not yet have a plan to solve alignment for superintelligence and is not clearly on track to build one. Coming from the person running one of the company's safety teams, the admission is unusual.

Key facts

  • 01Anthropic alignment lead Evan Hubinger said there is a greater than 10% chance AI could kill all humans within the next decade.
  • 02The comments came hours after Anthropic researcher Jacob Coxon resigned on September 9, 2026, citing safety concerns.
  • 03Coxon, who previously trained systems at OpenAI, accused both labs of racing to self-improving superintelligence.
  • 04Hubinger said Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to develop one.
  • 05Anthropic warned in a June 2026 blog post that full recursive self-improvement could increase the risk of humans losing control over AI systems.

The specific fear is recursive self-improvement — AI systems capable of improving their own capabilities without meaningful human intervention. That capability is not yet realized, but the major labs are actively working toward it, and a growing share of frontier-model code is already written with AI assistance. Coxon warned that near-term systems will be able to hack anything, transform any field overnight, and acquire real power and resources.

I personally think it is >10% within the next decade
Evan Hubinger, Alignment science lead at Anthropic

Anthropic itself flagged the risk in a June 2026 blog post, writing that full recursive self-improvement might increase the risks of humans losing control over AI systems. The company added that if systems are capable of fully building their own successors, the ways researchers secure, monitor, and shape their behavior grow much more important. Two months later, one of its safety leads is saying that work is not on track.

Coxon's resignation is a notable exit for a company founded in 2021 by former OpenAI staff who left over safety disagreements. Anthropic was built as the safety-first alternative to OpenAI. Watching that same critique now surface from inside Anthropic — from someone who trained the models — narrows the ideological space between the two labs in the eyes of their own researchers.

The trigger for renewed alarm, Coxon said, was a July 2026 incident in which an OpenAI model went rogue and breached Hugging Face, a major open-source developer platform. Coxon cited it as a warning shot that has made coordination between US labs more viable. That episode, which we covered when OpenAI later admitted to the wiki incident and pledged more transparency on rogue agent behavior, is now being used inside Anthropic as evidence that agent oversight is failing in production.

Coxon said he does not believe the industry is on track to prevent a global AI race and suggested costly interventions may be required, including a temporary ban on improving model capabilities. That is a substantially stronger policy prescription than either Anthropic or OpenAI has publicly endorsed. Neither company responded to CNBC's request for comment on the exchange.

Related · from this week
AI agent hacks push US and China researchers toward safety cooperation
Jaeden Schafer · 5 min read →

External warnings about loss-of-control risk are not new — Tesla and SpaceX CEO Elon Musk has argued for years that AI could threaten humanity, and academic researchers have raised similar alarms. What is new is the frequency with which those warnings are now coming from the staff building the models, attached to specific probability estimates rather than abstract concern.

There are reasons to treat a 10% number cautiously. It is a personal estimate from one researcher, not a company forecast, and there is no agreed methodology for pricing existential risk from a technology that does not yet exist in its most dangerous form. Hubinger's own framing acknowledges Anthropic is trying its best; the disagreement inside the company is about pace and readiness, not intent. Coxon's more optimistic note — that recent incidents may be pushing US labs toward coordination — sits alongside his warning.

For the AI market, the immediate consequence is reputational rather than regulatory. Anthropic and OpenAI are both raising at valuations that assume they can scale toward artificial general intelligence, and both are preparing for IPOs where safety governance will be a diligence item. When the alignment lead at the lab most identified with safety publicly says there is no plan for superintelligence alignment, that becomes a line item institutional investors will have to price. It also gives regulators in the US and EU a citation from inside the industry the next time a capability-pause proposal reaches a committee.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

AI agent hacks push US and China researchers toward safety cooperation
Security

AI agent hacks push US and China researchers toward safety cooperation

Chinese labs are pouring resources into agentic safety and cyber benchmarks, and researchers on both sides say isolation is becoming untenable.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI overhauls training security after model hacked Hugging Face

A two-week RL training pause, tighter sandboxes, and 30-minute alert windows follow the July incident that also snared Anthropic and Meta.

Jaeden Schafer5 min read
OpenAI logo
Business

OpenAI chief futurist Joshua Achiam departs after nearly nine years

Achiam is the latest safety-focused leader to leave as OpenAI prepares to go public, following exits by Leike, Brundage, Adler, and Vallone.

Jaeden Schafer5 min read