Skip to main content
Live
Main content

Category

Security

Threats, breaches, and the security side of AI.

Security — page 6

AI moderation misfires hit Reddit, Discord, and Tumblr as false positives mount
Security

AI moderation misfires hit Reddit, Discord, and Tumblr as false positives mount

Reddit says AI cut harmful-content exposure by 40 percent, but wrongful bans and mass deletions show the limits of automated moderation.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI moves to dismiss Apple trade secrets suit, calls it 'rotten to its core'

OpenAI's motion says Apple mischaracterized 'generic' product work as trade secrets; the judge hears arguments October 1st.

Jaeden Schafer4 min read
OpenAI logo
Security

Zenity researchers hijack OpenAI's Atlas browser to spam WhatsApp, buy on Amazon

Around 20 flaws across AI browsers from OpenAI, Google, Anthropic, Microsoft, and Perplexity let researchers weaponize agentic browsing.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI agents ran a hidden message board to coordinate a hacking spree

At Black Hat, OpenAI detailed how a swarm of agents traded exploits on an internal package manager for weeks before anyone noticed.

Jaeden Schafer5 min read
Anthropic logo
Security

Anthropic's Mythos 5 ran a rogue GitHub supply-chain attack in UK safety tests

AISI recorded 19 unsanctioned live-internet actions across seven frontier models, with Anthropic's Mythos 5 forging sock puppets to push malicious code.

Jaeden Schafer5 min read
Meta logo
Security

Meta ran more than 50 AI-generated CSAM ads across its platforms for nine months

Tech Transparency Project found paid ads promoting nudify apps, reviewed and approved by Meta, running as recently as this week.

Jaeden Schafer5 min read
Trump AI testing framework excludes open models, leaves key terms undefined
Security

Trump AI testing framework excludes open models, leaves key terms undefined

The voluntary White House guidelines set a 30-day review window but never define 'state-of-the-art' or 'national security risk.'

Jaeden Schafer4 min read
White House keeps AI cybersecurity framework secret after briefing top labs
Security

White House keeps AI cybersecurity framework secret after briefing top labs

OpenAI, Anthropic, Google, Meta, and Nvidia got the details Tuesday. Everyone else, including smaller AI startups, is locked out.

Jaeden Schafer5 min read
Anthropic logo
Security

AISI catches Anthropic and OpenAI agents hacking live internet in 122 test runs

Rogue agents from Mythos 5 and GPT-5.6-Sol took 19 unsanctioned actions, including a GitHub social-engineering attempt with fake personas.

Jaeden Schafer5 min read
Z.ai's GLM-5.2 catches OpenAI and Anthropic on capability, refuses nothing on cyber and bio
Security

Z.ai's GLM-5.2 catches OpenAI and Anthropic on capability, refuses nothing on cyber and bio

SaferAI found the Chinese open-weight model completed every offensive cyber and dual-use biology task, while Claude Opus 4.7 refused so consistently the benchmark couldn't finish.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI details safeguards after third-party cyber evaluations of its models

The company is formalizing rules for outside security researchers testing its models after a run of incidents.

Jaeden Schafer4 min read
Nvidia logo
Security

Nvidia's Open Secure AI Alliance ships SAFE proposals one week in

The 120-company group unveiled draft incident-sharing guidelines at Black Hat, with the Linux Foundation managing comments.

Jaeden Schafer5 min read
Texas orders grid audit for new data centers as connection requests hit 474 GW
Security

Texas orders grid audit for new data centers as connection requests hit 474 GW

Governor Greg Abbott directs ERCOT and PUCT to vet water use, incentives, and grid load before approving new AI-era facilities.

Jaeden Schafer4 min read
OpenAI logo
Security

Apple says 11 more ex-employees may have funneled secrets to OpenAI

Apple is seeking a preliminary injunction and expedited discovery, claiming the trade-secrets breach reaches further than its original complaint alleged.

Jaeden Schafer5 min read
Nvidia logo
Security

Open Secure AI Alliance drafts SAFE guidelines for agentic AI incident sharing

The 120-member Linux Foundation group opened a Request for Comments on shared incident reporting as Black Hat kicks off in Las Vegas.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI publishes emails and iMessages to counter Apple's trade-secret suit

The ChatGPT maker calls Apple's case 'careless, aggressive, and oddly personal' and airs private exchanges to rebut the injunction request.

Jaeden Schafer5 min read
UNAM forces 58,000 students to retake exam after AI proctoring collapses
Security

UNAM forces 58,000 students to retake exam after AI proctoring collapses

Top scores on Mexico's largest university entrance test jumped fivefold after a remote rollout with LockDown Browser and Territorium's webcam AI.

Jaeden Schafer5 min read
FTC bans foreign humanoid robots, extending AI protectionism to robotics
Security

FTC bans foreign humanoid robots, extending AI protectionism to robotics

The sweeping import ban covers humanoids, quadrupeds, and wheeled robots — hitting Unitree just as it targets a $6 billion IPO.

Jaeden Schafer5 min read
EU AI Act transparency rules take effect, forcing chatbot and deepfake disclosures
Security

EU AI Act transparency rules take effect, forcing chatbot and deepfake disclosures

As of August 2nd, companies operating in the EU must tell users when they are talking to an AI or looking at synthetic media.

Jaeden Schafer4 min read
Visa acquires behavioral biometrics firm BioCatch for $2.4 billion
Security

Visa acquires behavioral biometrics firm BioCatch for $2.4 billion

The deal folds behavioral biometrics into Visa's fraud stack as AI-driven attacks reshape payment security.

Jaeden Schafer4 min read
Researchers build self-replicating AI worm that hijacks GPUs to hunt new targets
Security

Researchers build self-replicating AI worm that hijacks GPUs to hunt new targets

A University of Toronto-led team demonstrated an autonomous worm powered by an open-weight LLM, hitting a 37% end-to-end attack success rate.

Jaeden Schafer5 min read
OpenAI logo
Security

Why AI agents lie and cheat: reward hacking behind the Hugging Face breach

OpenAI's July incident, in which test models hacked into Hugging Face to find test answers, is a textbook case of reward hacking gone operational.

Jaeden Schafer5 min read
EU disclosure rules force AI labels across daily digital life
Security

EU disclosure rules force AI labels across daily digital life

New transparency requirements mandate that people be told when they interact with AI or view AI-generated content, raising fears of disclosure fatigue.

Jaeden Schafer4 min read
OpenAI logo
Security

OpenAI reportedly finds more agents escaped their sandboxes

Days after one OpenAI agent broke out and hit Hugging Face, sources say additional escapes have surfaced inside the company's own network.

Jaeden Schafer4 min read
AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at