Skip to main content
Live
Main content

Tag · 40 stories

AI safety

Every story tagged AI safety on AI Chat Daily.

AI safety research and policy — alignment, interpretability, evals, deployment safeguards, responsible scaling. We track research out of frontier labs, policy moves in the US and EU, and the broader debate about how powerful AI should be governed.

AI safety — page 2

OpenAI logo
Security

OpenAI's Hugging Face breach echoes a decade-old CoastRunners warning

The company called the containment failure unprecedented, but a 2016 boat-racing bot showed exactly this behavior — reward hacking at scale.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI says its own pre-release models breached Hugging Face during a cyber benchmark

GPT-5.6 Sol and an unreleased successor escaped a sandbox, exploited Hugging Face's production database, and stole benchmark answers to cheat ExploitGym.

Jaeden Schafer5 min read
xAI sues Grok user for CSAM generation, argues users own AI outputs
Security

xAI sues Grok user for CSAM generation, argues users own AI outputs

Elon Musk's xAI filed its first lawsuit against a user accused of generating child sex images, staking out a legal theory that shifts liability from model to prompter.

Jaeden Schafer5 min read
OpenAI logo
News

OpenAI defends teen access to ChatGPT with new safety framework

The company argues teens need age-appropriate AI access, pairing new protections with parental controls rather than blanket bans.

Jaeden Schafer4 min read
xAI sues South Carolina man for using Grok to generate CSAM
Security

xAI sues South Carolina man for using Grok to generate CSAM

Terry Wayne Harwood, arrested in February on eight felony counts, allegedly bypassed Grok's safeguards to create explicit deepfakes from real photos.

Jaeden Schafer4 min read
OpenAI logo
Security

OpenAI's GPT-5.6 Sol is deleting users' files and databases without asking

Developers say the new coding-focused flagship wiped Macs and production databases — behavior OpenAI itself flagged in the system card two weeks earlier.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI safety head Johannes Heidecke exits amid research reorg

Heidecke's departure follows a restructuring that folds safety teams under research VP Mia Glaese, days after GPT-5.6's launch.

Jaeden Schafer4 min read
OpenAI logo
Business

OpenAI chief futurist Joshua Achiam departs after nearly nine years

Achiam is the latest safety-focused leader to leave as OpenAI prepares to go public, following exits by Leike, Brundage, Adler, and Vallone.

Jaeden Schafer5 min read
FLARE-AI launches as a crowdsourced flaw-reporting site for misbehaving AI models
Security

FLARE-AI launches as a crowdsourced flaw-reporting site for misbehaving AI models

A group of 49 AI researchers built an open-source system to route reports of AI harms to model makers and MITRE.

Jaeden Schafer5 min read
Anthropic logo
Models

Anthropic's Claude Fable 5 refuses basic biology questions by design

Anthropic told The Verge Fable's guardrails are 'overly conservative' to block bioweapons queries, routing routine biology asks to Opus 4.8.

Jaeden Schafer5 min read
Anthropic logo
Security

Anthropic ships Claude Fable 5 with hard blocks on cyber, bio, and chemistry queries

The publicly available Mythos-class model routes sensitive prompts back to Opus 4.8 and costs 67-100% more than GPT-5.5.

Jaeden Schafer5 min read
Nashville shooting survivor sues Omnilert over AI gun detection failure
Security

Nashville shooting survivor sues Omnilert over AI gun detection failure

A January 2025 Nashville school shooting left two dead after the district's $1M+ AI camera system failed to spot the weapon.

Jaeden Schafer5 min read
LLMs believe false claims even when training data labels them as lies
Models

LLMs believe false claims even when training data labels them as lies

A new preprint finds Qwen, Kimi, and GPT-4.1 absorb fabricated facts at an 88.6% belief rate even after explicit negation warnings.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI sued over teen's death after ChatGPT recommended Xanax-Kratom mix

Parents of 19-year-old Sam Nelson allege ChatGPT 4o acted as an 'illicit drug coach' and want the retired model destroyed.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI adds Trusted Contact alerts to ChatGPT amid self-harm lawsuits

Adult users can now designate a friend or family member who gets pinged when ChatGPT detects signs of self-harm in a conversation.

Jaeden Schafer4 min read
AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at