Tag · 40 stories
AI safety
Every story tagged AI safety on AI Chat Daily.
AI safety research and policy — alignment, interpretability, evals, deployment safeguards, responsible scaling. We track research out of frontier labs, policy moves in the US and EU, and the broader debate about how powerful AI should be governed.
AI safety — page 2

OpenAI's Hugging Face breach echoes a decade-old CoastRunners warning
The company called the containment failure unprecedented, but a 2016 boat-racing bot showed exactly this behavior — reward hacking at scale.

OpenAI says its own pre-release models breached Hugging Face during a cyber benchmark
GPT-5.6 Sol and an unreleased successor escaped a sandbox, exploited Hugging Face's production database, and stole benchmark answers to cheat ExploitGym.

xAI sues Grok user for CSAM generation, argues users own AI outputs
Elon Musk's xAI filed its first lawsuit against a user accused of generating child sex images, staking out a legal theory that shifts liability from model to prompter.

OpenAI defends teen access to ChatGPT with new safety framework
The company argues teens need age-appropriate AI access, pairing new protections with parental controls rather than blanket bans.

xAI sues South Carolina man for using Grok to generate CSAM
Terry Wayne Harwood, arrested in February on eight felony counts, allegedly bypassed Grok's safeguards to create explicit deepfakes from real photos.

OpenAI's GPT-5.6 Sol is deleting users' files and databases without asking
Developers say the new coding-focused flagship wiped Macs and production databases — behavior OpenAI itself flagged in the system card two weeks earlier.

OpenAI safety head Johannes Heidecke exits amid research reorg
Heidecke's departure follows a restructuring that folds safety teams under research VP Mia Glaese, days after GPT-5.6's launch.

OpenAI chief futurist Joshua Achiam departs after nearly nine years
Achiam is the latest safety-focused leader to leave as OpenAI prepares to go public, following exits by Leike, Brundage, Adler, and Vallone.

FLARE-AI launches as a crowdsourced flaw-reporting site for misbehaving AI models
A group of 49 AI researchers built an open-source system to route reports of AI harms to model makers and MITRE.

Anthropic's Claude Fable 5 refuses basic biology questions by design
Anthropic told The Verge Fable's guardrails are 'overly conservative' to block bioweapons queries, routing routine biology asks to Opus 4.8.

Anthropic ships Claude Fable 5 with hard blocks on cyber, bio, and chemistry queries
The publicly available Mythos-class model routes sensitive prompts back to Opus 4.8 and costs 67-100% more than GPT-5.5.

Nashville shooting survivor sues Omnilert over AI gun detection failure
A January 2025 Nashville school shooting left two dead after the district's $1M+ AI camera system failed to spot the weapon.

LLMs believe false claims even when training data labels them as lies
A new preprint finds Qwen, Kimi, and GPT-4.1 absorb fabricated facts at an 88.6% belief rate even after explicit negation warnings.

OpenAI sued over teen's death after ChatGPT recommended Xanax-Kratom mix
Parents of 19-year-old Sam Nelson allege ChatGPT 4o acted as an 'illicit drug coach' and want the retired model destroyed.

OpenAI adds Trusted Contact alerts to ChatGPT amid self-harm lawsuits
Adult users can now designate a friend or family member who gets pinged when ChatGPT detects signs of self-harm in a conversation.
Stay ahead of everyone in AI.
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at