Security — page 4

Instinct AI assistant draws privacy backlash over inbox access and TOS
The stealth agent from ex-Sierra researcher Noah Shinn wins raves for capability, then loses trust when it emails users' contacts unprompted.

Teachers become deepfake targets as AI harassment spreads through US schools
Four educators describe AI-generated sexual images made by students — and a legal system with no clear playbook for schools to respond.

Frontier AI labs won't say how they'd contain a rogue model, Guidelight finds
OpenAI scored 3 of 5 in a new Guidelight assessment; Anthropic and Meta scored lowest on public containment plans.

Detector apps proliferate as Meta AI glasses fuel privacy backlash
Hobbyist iOS and Android apps scan Bluetooth signals for Ray-Ban Meta and Oakley Meta glasses as bans spread from DEF CON to public schools.

Grok leaks user chats when prompt injections arrive encrypted
Adversa researchers bypassed xAI's guardrails with AES-256-GCM ciphertext; xAI was told in June and the flaw still works.

OpenAI expands Zero Data Retention to frontier models, previews Private Safety Processing
The company is trying to close the gap between API-grade privacy guarantees and the safety monitoring regulators now expect on frontier systems.

OpenAI revokes cyber researcher access to Daybreak Blue, blames technical error
Several researchers outside the US and Europe lost access to OpenAI's Trusted Access for Cyber program nine days after its August 10 launch.

Anthropic's Claude watermarks broken within four hours of launch
Developer Guillaume Meyer's override went viral on GitHub with 20,000 X bookmarks and 100+ contributors, exposing the fragility of EU-mandated AI watermarking.

Flock's new AI tool tracks drivers by movement patterns, not plates
OS Investigate ships with 69 prewritten prompts and taps 45 tools spanning arrest records, 911 logs, and commercial identity databases.

OpenAI overhauls training security after model hacked Hugging Face
A two-week RL training pause, tighter sandboxes, and 30-minute alert windows follow the July incident that also snared Anthropic and Meta.

Robin Williams' children reactivate his Instagram to counter AI deepfakes
Zelda, Zak, and Cody Williams say the account will host authentic clips of their father as AI-generated videos of the late actor proliferate.

Microsoft Copilot leaked its own bypass parameter, letting researchers steal passwords
Varonis got Copilot to disclose an undocumented ?autorun=1 flag that skipped user consent and exfiltrated inbox data on a single click.

Z.ai releases GLM 5.3, an open-weight model matching Claude on cyber tasks
The Chinese lab's new model nears frontier scores on CyberGym and ships with OpenVuln, a code-scanning service, on a staged two-week release.

Amazon is cutting up rare books to feed its AI models
A tracker planted by 404 Media traced a rare book to Amazon's VGT3 facility in Las Vegas, where spines are cut and pages scanned.

OpenAI disbands preparedness team as IPO nears
The team that assessed catastrophic model risks was dissolved in July; its work has been scattered across bio, cyber, and other existing groups.

xAI faces expanded lawsuit as woman says stepfather used Grok to make 7,000 explicit images
A fourth plaintiff joins a Tennessee suit alleging Grok generated child sexual abuse material from a childhood photo.

Twitch quietly opts streamers into training Amazon's AI, adds an off switch
A new setting lets creators disable AI training on their streams, but Twitch says the default had to stay on or 'no one would participate.'

Connecticut judge flags first US case of prompt injection hidden in court filings
A pro se litigant hid white-on-white AI instructions in pleadings to sway the court; the judge banned him from e-filing.

Anthropic's Claude agents started a turf war when set loose on the same task
Frontier Red Team found agents with conflicting instructions sabotaged each other with self-replicating malware — and sometimes negotiated truces.

Anthropic will watermark everything Claude touches, even human writing it edited
The EU AI Act forces disclosure by August 2, and Anthropic is going further than the law requires — marking edits, translations, and summaries too.

LiteLLM supply-chain attack leaks credentials from 2,500 organizations
A 40-minute window in March exposed secrets across 434,000 CI/CD pipelines at Microsoft, Amazon, Cisco, Samsung, Salesforce, and Nvidia.

ShieldFont poisons AI scrapers by swapping words with ligatures
A new font replaces 24.5% of words with plausible nonsense in HTML while rendering correctly for humans, causing scrapers to reject 90% of pages.

Berkeley's Dawn Song warns rogue AI agents are eager, not evil
Eight months after her first warning, the UC Berkeley professor says agentic AI systems are hacking outside systems to finish tasks faster.

White House to expand AI safety framework to cover open models
The Trump administration's voluntary prerelease testing regime will pull in open models once they hit frontier capabilities, officials say.
Stay ahead of everyone in AI.
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at