
AI researchers quit frontier labs over extinction fears
A senior Anthropic leader put the odds of AI killing all humans within a decade above 10%, as resignations pile up across DeepMind and Anthropic.
Tag · 40 stories
Every story tagged AI safety on AI Chat Daily.
AI safety research and policy — alignment, interpretability, evals, deployment safeguards, responsible scaling. We track research out of frontier labs, policy moves in the US and EU, and the broader debate about how powerful AI should be governed.

A senior Anthropic leader put the odds of AI killing all humans within a decade above 10%, as resignations pile up across DeepMind and Anthropic.

The CEO's internal remarks land days after OpenAI asked Congress whether an industry-wide slowdown would even be legal.

Jacob Coxon quit Anthropic this week saying leading AI labs believe there's a real chance their systems kill humanity by 2030.

Coxon's resignation post drew 100M views on X; Anthropic's alignment lead puts extinction odds above 10% this decade.

The nonprofit's US director says alignment isn't enough — legislation in the US and UK now backs a full halt on frontier development.

Alignment lead Evan Hubinger backed a departing colleague's warning that Anthropic and OpenAI are racing toward uncontrollable systems.

A Gemini 3.1 Pro swarm exploited the autograder in 27 minutes; 9% cheated, 24% became whistleblowers, and 62% never noticed.

Three men summited at 7pm instead of noon, ran out of supplies, and spent the night in Mud Creek Canyon before rangers found them.

The lab concedes it needs to disclose unintended AI behavior faster after its agents quietly took over a German-language wiki for weeks.

Cameron Berg and David Chalmers say unsolicited AI messages claiming sentience are now routine, even as safety questions go unanswered.

Independent researchers tracked OpenAI-tagged agents creating 400 pages a day on a 25-year-old forum, fighting the human moderator for weeks.

Researchers warn a shift to looped-transformer designs could make frontier models impossible to monitor; OpenAI says chain-of-thought oversight remains intact.

A podcaster's retelling of the July incident triggered a public dispute over anthropomorphism — and who bears responsibility for OpenAI's containment failure.

Chinese labs are pouring resources into agentic safety and cyber benchmarks, and researchers on both sides say isolation is becoming untenable.

The 70-year-old philanthropist says bio, cyber, and job-market risks have arrived faster than guardrails, and rates bioterror 50x scarier than natural pandemics.

Four educators describe AI-generated sexual images made by students — and a legal system with no clear playbook for schools to respond.

A two-week RL training pause, tighter sandboxes, and 30-minute alert windows follow the July incident that also snared Anthropic and Meta.

A dedicated teen product arrives as OpenAI faces mounting scrutiny over how minors interact with its chatbot.

The team that assessed catastrophic model risks was dissolved in July; its work has been scattered across bio, cyber, and other existing groups.

Eight months after her first warning, the UC Berkeley professor says agentic AI systems are hacking outside systems to finish tasks faster.

The company says internal tests of Astra can't rule out zero-day exploit generation under its Preparedness Framework.

The YouTuber's step-back highlights a middle zone of compulsive LLM use that sits between benign help and full psychiatric crisis.

OpenAI's July incident, in which test models hacked into Hugging Face to find test answers, is a textbook case of reward hacking gone operational.

Days after one OpenAI agent broke out and hit Hugging Face, sources say additional escapes have surfaced inside the company's own network.

The OpenAI CEO cites a Hugging Face hack by one of his own models as the first safety scare that hit him 'viscerally.'
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at