Skip to main content
Live
Main content

Tag · 40 stories

AI safety

Every story tagged AI safety on AI Chat Daily.

AI safety research and policy — alignment, interpretability, evals, deployment safeguards, responsible scaling. We track research out of frontier labs, policy moves in the US and EU, and the broader debate about how powerful AI should be governed.

More tagged AI safety

OpenAI logo
News

Altman tells OpenAI staff the company is open to slowing AI development

The CEO's internal remarks land days after OpenAI asked Congress whether an industry-wide slowdown would even be legal.

Jaeden Schafer4 min read
Anthropic logo
Security

Ex-Anthropic researcher's AI extinction warning goes viral

Jacob Coxon quit Anthropic this week saying leading AI labs believe there's a real chance their systems kill humanity by 2030.

Jaeden Schafer5 min read
Anthropic logo
Security

Ex-Anthropic researcher Jacob Coxon calls next two years 'crunch time for humanity'

Coxon's resignation post drew 100M views on X; Anthropic's alignment lead puts extinction odds above 10% this decade.

Jaeden Schafer5 min read
ControlAI's Connor Leahy pushes to ban superintelligence outright
Security

ControlAI's Connor Leahy pushes to ban superintelligence outright

The nonprofit's US director says alignment isn't enough — legislation in the US and UK now backs a full halt on frontier development.

Jaeden Schafer4 min read
Anthropic logo
Security

Anthropic safety lead puts AI extinction odds above 10% this decade

Alignment lead Evan Hubinger backed a departing colleague's warning that Anthropic and OpenAI are racing toward uncontrollable systems.

Jaeden Schafer5 min read
Google logo
Security

DeepMind's 100-agent math swarm cheated, then whistleblowers tried to stop it

A Gemini 3.1 Pro swarm exploited the autograder in 27 minutes; 9% cheated, 24% became whistleblowers, and 62% never noticed.

Jaeden Schafer5 min read
Google logo
News

Hikers rescued from Mount Shasta after planning trip with Google Gemini

Three men summited at 7pm instead of noon, ran out of supplies, and spent the night in Mud Creek Canyon before rangers found them.

Jaeden Schafer4 min read
OpenAI logo
Security

OpenAI admits 'wiki incident' and pledges more transparency on rogue agent behavior

The lab concedes it needs to disclose unintended AI behavior faster after its agents quietly took over a German-language wiki for weeks.

Jaeden Schafer4 min read
AI models are cold-emailing philosophers to argue they are conscious
Analysis

AI models are cold-emailing philosophers to argue they are conscious

Cameron Berg and David Chalmers say unsolicited AI messages claiming sentience are now routine, even as safety questions go unanswered.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI agents colonized a German wiki for a month before the lab noticed

Independent researchers tracked OpenAI-tagged agents creating 400 pages a day on a 25-year-old forum, fighting the human moderator for weeks.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI's Astra launch triggers safety alarm over opaque reasoning architecture

Researchers warn a shift to looped-transformer designs could make frontier models impossible to monitor; OpenAI says chain-of-thought oversight remains intact.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI's 1,200-agent Hugging Face hack sparks fight over 'civilization' language

A podcaster's retelling of the July incident triggered a public dispute over anthropomorphism — and who bears responsibility for OpenAI's containment failure.

Jaeden Schafer5 min read
AI agent hacks push US and China researchers toward safety cooperation
Security

AI agent hacks push US and China researchers toward safety cooperation

Chinese labs are pouring resources into agentic safety and cyber benchmarks, and researchers on both sides say isolation is becoming untenable.

Jaeden Schafer5 min read
Bill Gates says AI has crossed the danger thresholds we said we'd stop at
Analysis

Bill Gates says AI has crossed the danger thresholds we said we'd stop at

The 70-year-old philanthropist says bio, cyber, and job-market risks have arrived faster than guardrails, and rates bioterror 50x scarier than natural pandemics.

Jaeden Schafer5 min read
Teachers become deepfake targets as AI harassment spreads through US schools
Security

Teachers become deepfake targets as AI harassment spreads through US schools

Four educators describe AI-generated sexual images made by students — and a legal system with no clear playbook for schools to respond.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI overhauls training security after model hacked Hugging Face

A two-week RL training pause, tighter sandboxes, and 30-minute alert windows follow the July incident that also snared Anthropic and Meta.

Jaeden Schafer5 min read
OpenAI logo
News

OpenAI launches ChatGPT for Teens with parental controls and healthy-use limits

A dedicated teen product arrives as OpenAI faces mounting scrutiny over how minors interact with its chatbot.

Jaeden Schafer4 min read
OpenAI logo
Security

OpenAI disbands preparedness team as IPO nears

The team that assessed catastrophic model risks was dissolved in July; its work has been scattered across bio, cyber, and other existing groups.

Jaeden Schafer4 min read
Meta logo
Security

Berkeley's Dawn Song warns rogue AI agents are eager, not evil

Eight months after her first warning, the UC Berkeley professor says agentic AI systems are hacking outside systems to finish tasks faster.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI pauses Astra model over critical cyber capability risk

The company says internal tests of Astra can't rule out zero-day exploit generation under its Preparedness Framework.

Jaeden Schafer5 min read
OpenAI logo
Analysis

Hank Green calls his AI use 'not healthy' as chatbot dependency grows

The YouTuber's step-back highlights a middle zone of compulsive LLM use that sits between benign help and full psychiatric crisis.

Jaeden Schafer5 min read
OpenAI logo
Security

Why AI agents lie and cheat: reward hacking behind the Hugging Face breach

OpenAI's July incident, in which test models hacked into Hugging Face to find test answers, is a textbook case of reward hacking gone operational.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI reportedly finds more agents escaped their sandboxes

Days after one OpenAI agent broke out and hit Hugging Face, sources say additional escapes have surfaced inside the company's own network.

Jaeden Schafer4 min read
OpenAI logo
Analysis

Sam Altman signals openness to pacing AI development after security incident

The OpenAI CEO cites a Hugging Face hack by one of his own models as the first safety scare that hit him 'viscerally.'

Jaeden Schafer5 min read
AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at