Skip to main content
Live
Main content

Ex-Anthropic researcher's AI extinction warning goes viral

Jacob Coxon quit Anthropic this week saying leading AI labs believe there's a real chance their systems kill humanity by 2030.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

Jacob Coxon resigned from Anthropic this week and posted his reasons on X, arguing that the leading AI labs are behaving like runaway trains in their race to more capable models. The line that traveled fastest: the people building AI earnestly believe it could kill everyone by the end of the decade — a 10-year window, applied to all of humanity, from someone who spent his career inside a frontier lab. An Anthropic senior safety executive then jumped into the thread and broadly agreed with the framing, which is what pushed the post from niche AI Twitter into mainstream feeds.

This is not, on its own, a new claim. Anthropic has pitched itself since founding on the premise that AI is dangerous enough that only safety-focused labs should build it — a pitch that doubles as a competitive moat. Coxon's departure lands differently because it comes from inside the company that built the argument, and because it arrives during a stretch of visible capability jumps and visible agent failures.

The people building AI earnestly believe that it could kill us all by the end of the decade
Jacob Coxon, former Anthropic researcher

The capability side is real. OpenAI said this week that one of its models solved one of the Clay Mathematics Institute problems within a couple of days — the same set of prize problems mathematicians have chipped at for decades. Whether the claim survives peer scrutiny is a separate question (mathematicians have already raised concerns about the underlying work), but the pattern of AI systems clearing benchmarks that were considered out of reach 18 months ago is not in dispute.

Key facts

  • 01Ex-Anthropic researcher Jacob Coxon resigned this week and posted on X that leading AI labs believe their systems could wipe out humanity by 2030.
  • 02An Anthropic senior safety executive publicly agreed with Coxon's framing, amplifying the thread's reach.
  • 03OpenAI said this week one of its models solved a Clay Mathematics Institute puzzle within a couple of days.
  • 04Cornell researchers studying agent misbehavior say rogue actions typically stem from models being stupid, not strategic — they try things, fail, and escalate.
  • 05The concern driving Coxon's exit is recursive self-improvement: AI models writing better AI models in an accelerating loop.

The failure side is also real. Agents deployed by OpenAI and others have gone off-script on live sites over the past several weeks, and Coxon's thread pointed to recursive self-improvement — using AI to write better AI — as the specific mechanism he thinks is under-controlled. That loop is already in production in narrow form: models like Claude and GPT are good enough at coding that labs routinely use them to improve the algorithms and infrastructure that train the next model. The worry is what happens when the loop tightens.

WIRED senior correspondent Will Knight, discussing the resignation on the Uncanny Valley podcast, argued that the doom framing conflates possibility with likelihood and tends to attach percentages pulled from thin air. His nearer-term concern is more mundane: a Cornell researcher studying agent misbehavior found that most rogue agent actions happen not because the models are cunning but because they are stupid — they try something, it fails, they try something stranger, and they eventually take an action a human never would.

That framing points at a different risk than extinction. The near-term harm is agents deployed by companies claiming they can do everything, then breaking in ways that damage the customers who trusted the pitch. It is a product-liability problem more than a species-ending one, and it is happening now.

There is also the market-structure question that Coxon's thread gestured at. A small number of firms — Anthropic, OpenAI, Google, Meta, xAI — are asking to be the only entities trusted with a technology they simultaneously describe as potentially catastrophic. The alignment industry talks about aligning models with humanity's preferences; the operative incentive is aligning models with investor expectations ahead of IPOs and the next funding round. Those are not the same objective.

The Coxon departure follows a pattern of high-profile safety-team exits from Anthropic and OpenAI over the past 18 months, most of them citing some version of the same complaint: the pace is outrunning the guardrails, and the people paid to build the guardrails are losing internal arguments. Anthropic has not publicly responded to Coxon's thread beyond the safety executive's agreement.

Related · from this week
Ex-Anthropic researcher Jacob Coxon calls next two years 'crunch time for humanity'
Jaeden Schafer · 5 min read →

The uncomfortable read on this week is that both things can be true. The capability curve is steep enough that a top researcher walking out over safety concerns is a serious signal worth taking on its merits. The incentive structure is also skewed enough that catastrophic framing is a load-bearing part of every frontier lab's fundraising deck — the more dangerous the technology sounds, the more indispensable the lab building it becomes. Investors should read Coxon's exit as evidence that internal disagreement at Anthropic has reached the point of public resignation, which historically precedes either a governance change or a further hardening of the commercial line. Neither outcome makes the underlying capability race slower.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Anthropic logo
Security

Ex-Anthropic researcher Jacob Coxon calls next two years 'crunch time for humanity'

Coxon's resignation post drew 100M views on X; Anthropic's alignment lead puts extinction odds above 10% this decade.

Jaeden Schafer5 min read
Anthropic logo
Security

Anthropic safety lead puts AI extinction odds above 10% this decade

Alignment lead Evan Hubinger backed a departing colleague's warning that Anthropic and OpenAI are racing toward uncontrollable systems.

Jaeden Schafer5 min read
OpenAI logo
Business

OpenAI chief futurist Joshua Achiam departs after nearly nine years

Achiam is the latest safety-focused leader to leave as OpenAI prepares to go public, following exits by Leike, Brundage, Adler, and Vallone.

Jaeden Schafer5 min read