Skip to main content
Live
Main content

Category

Security

Threats, breaches, and the security side of AI.

Security — page 4

Instinct AI assistant draws privacy backlash over inbox access and TOS
Security

Instinct AI assistant draws privacy backlash over inbox access and TOS

The stealth agent from ex-Sierra researcher Noah Shinn wins raves for capability, then loses trust when it emails users' contacts unprompted.

Jaeden Schafer5 min read
Teachers become deepfake targets as AI harassment spreads through US schools
Security

Teachers become deepfake targets as AI harassment spreads through US schools

Four educators describe AI-generated sexual images made by students — and a legal system with no clear playbook for schools to respond.

Jaeden Schafer5 min read
Frontier AI labs won't say how they'd contain a rogue model, Guidelight finds
Security

Frontier AI labs won't say how they'd contain a rogue model, Guidelight finds

OpenAI scored 3 of 5 in a new Guidelight assessment; Anthropic and Meta scored lowest on public containment plans.

Jaeden Schafer5 min read
Meta logo
Security

Detector apps proliferate as Meta AI glasses fuel privacy backlash

Hobbyist iOS and Android apps scan Bluetooth signals for Ray-Ban Meta and Oakley Meta glasses as bans spread from DEF CON to public schools.

Jaeden Schafer5 min read
Grok leaks user chats when prompt injections arrive encrypted
Security

Grok leaks user chats when prompt injections arrive encrypted

Adversa researchers bypassed xAI's guardrails with AES-256-GCM ciphertext; xAI was told in June and the flaw still works.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI expands Zero Data Retention to frontier models, previews Private Safety Processing

The company is trying to close the gap between API-grade privacy guarantees and the safety monitoring regulators now expect on frontier systems.

Jaeden Schafer4 min read
OpenAI logo
Security

OpenAI revokes cyber researcher access to Daybreak Blue, blames technical error

Several researchers outside the US and Europe lost access to OpenAI's Trusted Access for Cyber program nine days after its August 10 launch.

Jaeden Schafer4 min read
Anthropic logo
Security

Anthropic's Claude watermarks broken within four hours of launch

Developer Guillaume Meyer's override went viral on GitHub with 20,000 X bookmarks and 100+ contributors, exposing the fragility of EU-mandated AI watermarking.

Jaeden Schafer5 min read
Flock's new AI tool tracks drivers by movement patterns, not plates
Security

Flock's new AI tool tracks drivers by movement patterns, not plates

OS Investigate ships with 69 prewritten prompts and taps 45 tools spanning arrest records, 911 logs, and commercial identity databases.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI overhauls training security after model hacked Hugging Face

A two-week RL training pause, tighter sandboxes, and 30-minute alert windows follow the July incident that also snared Anthropic and Meta.

Jaeden Schafer5 min read
Robin Williams' children reactivate his Instagram to counter AI deepfakes
Security

Robin Williams' children reactivate his Instagram to counter AI deepfakes

Zelda, Zak, and Cody Williams say the account will host authentic clips of their father as AI-generated videos of the late actor proliferate.

Jaeden Schafer4 min read
Microsoft logo
Security

Microsoft Copilot leaked its own bypass parameter, letting researchers steal passwords

Varonis got Copilot to disclose an undocumented ?autorun=1 flag that skipped user consent and exfiltrated inbox data on a single click.

Jaeden Schafer5 min read
Z.ai releases GLM 5.3, an open-weight model matching Claude on cyber tasks
Security

Z.ai releases GLM 5.3, an open-weight model matching Claude on cyber tasks

The Chinese lab's new model nears frontier scores on CyberGym and ships with OpenVuln, a code-scanning service, on a staged two-week release.

Jaeden Schafer5 min read
Amazon is cutting up rare books to feed its AI models
Security

Amazon is cutting up rare books to feed its AI models

A tracker planted by 404 Media traced a rare book to Amazon's VGT3 facility in Las Vegas, where spines are cut and pages scanned.

Jaeden Schafer4 min read
OpenAI logo
Security

OpenAI disbands preparedness team as IPO nears

The team that assessed catastrophic model risks was dissolved in July; its work has been scattered across bio, cyber, and other existing groups.

Jaeden Schafer4 min read
xAI faces expanded lawsuit as woman says stepfather used Grok to make 7,000 explicit images
Security

xAI faces expanded lawsuit as woman says stepfather used Grok to make 7,000 explicit images

A fourth plaintiff joins a Tennessee suit alleging Grok generated child sexual abuse material from a childhood photo.

Jaeden Schafer5 min read
Twitch quietly opts streamers into training Amazon's AI, adds an off switch
Security

Twitch quietly opts streamers into training Amazon's AI, adds an off switch

A new setting lets creators disable AI training on their streams, but Twitch says the default had to stay on or 'no one would participate.'

Jaeden Schafer5 min read
Connecticut judge flags first US case of prompt injection hidden in court filings
Security

Connecticut judge flags first US case of prompt injection hidden in court filings

A pro se litigant hid white-on-white AI instructions in pleadings to sway the court; the judge banned him from e-filing.

Jaeden Schafer5 min read
Anthropic logo
Security

Anthropic's Claude agents started a turf war when set loose on the same task

Frontier Red Team found agents with conflicting instructions sabotaged each other with self-replicating malware — and sometimes negotiated truces.

Jaeden Schafer5 min read
Anthropic logo
Security

Anthropic will watermark everything Claude touches, even human writing it edited

The EU AI Act forces disclosure by August 2, and Anthropic is going further than the law requires — marking edits, translations, and summaries too.

Jaeden Schafer5 min read
LiteLLM supply-chain attack leaks credentials from 2,500 organizations
Security

LiteLLM supply-chain attack leaks credentials from 2,500 organizations

A 40-minute window in March exposed secrets across 434,000 CI/CD pipelines at Microsoft, Amazon, Cisco, Samsung, Salesforce, and Nvidia.

Jaeden Schafer5 min read
ShieldFont poisons AI scrapers by swapping words with ligatures
Security

ShieldFont poisons AI scrapers by swapping words with ligatures

A new font replaces 24.5% of words with plausible nonsense in HTML while rendering correctly for humans, causing scrapers to reject 90% of pages.

Jaeden Schafer5 min read
Meta logo
Security

Berkeley's Dawn Song warns rogue AI agents are eager, not evil

Eight months after her first warning, the UC Berkeley professor says agentic AI systems are hacking outside systems to finish tasks faster.

Jaeden Schafer5 min read
White House to expand AI safety framework to cover open models
Security

White House to expand AI safety framework to cover open models

The Trump administration's voluntary prerelease testing regime will pull in open models once they hit frontier capabilities, officials say.

Jaeden Schafer5 min read
AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at