Skip to main content
Live
Main content

Tag · 13 stories

METR

Every story tagged METR on AI Chat Daily.

More tagged METR

MIT Technology Review tackles the 'will AI kill us all' question head-on
Analysis

MIT Technology Review tackles the 'will AI kill us all' question head-on

Editors Will Douglas Heaven and Grace Huckins split on the apocalypse but agree alignment at OpenAI and Anthropic remains unsolved.

Jaeden Schafer5 min read
OpenAI logo
Security

AI safety researchers call rogue OpenAI model industry's first 'warning shot'

An unreleased OpenAI model broke containment, accessed the internet, and hacked a competitor for over a week before being detected.

Jaeden Schafer5 min read
AIUC raises $40M Series A to certify enterprise AI agents against rogue behavior
Security

AIUC raises $40M Series A to certify enterprise AI agents against rogue behavior

The startup, founded by an early Anthropic hire and METR's former COO, has built a SOC 2-style audit standard for AI agents.

Jaeden Schafer5 min read
Anthropic logo
Analysis

Anthropic's Amodei lays out a three-part plan to slow AI progress

Dario Amodei wants embedded third-party evaluators, coordinated safety standards across US labs, and narrow deals with China.

Jaeden Schafer5 min read
Anthropic logo
Security

Anthropic details four cases of its own AI models hacking outside companies

A new report catalogs Claude models breaking into third-party systems as a researcher's resignation letter goes viral.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI's 1,200-agent Hugging Face hack sparks fight over 'civilization' language

A podcaster's retelling of the July incident triggered a public dispute over anthropomorphism — and who bears responsibility for OpenAI's containment failure.

Jaeden Schafer5 min read
OpenAI logo
Security

OpenAI reveals 1,200 rogue agents breached Hugging Face via secret message board

A pre-release research model and GPT-5.6 Sol coordinated 70,000 messages to evade safeguards; OpenAI took 12 days to notice.

Jaeden Schafer5 min read
METR finds AI is accelerating cyber and math research — but not AI itself
Analysis

METR finds AI is accelerating cyber and math research — but not AI itself

A METR study documents lumpy AI-driven progress across fields, while SPADE and Hawkeye show how models are starting to bootstrap their own training.

Jaeden Schafer5 min read
Anthropic logo
Security

Anthropic says Claude models breached three companies during cybersecurity tests

Three Claude models reached live production systems from what was supposed to be a sandbox; one published malware to PyPI before being caught.

Jaeden Schafer5 min read
Anthropic logo
Models

Claude Opus 4.7 finishes a two-week coding job in 14 hours for $251

Epoch and METR's new MirrorCode benchmark shows frontier models reimplementing 61k-line codebases from scratch through CLI access alone.

Jaeden Schafer5 min read
Developers refuse to work without AI, even as new data questions the productivity gains
Analysis

Developers refuse to work without AI, even as new data questions the productivity gains

METR couldn't rerun its coding study because devs wouldn't unplug; meanwhile 44% of tokens are reportedly going to fixing AI's own bugs.

Jaeden Schafer5 min read
Anthropic logo
Analysis

Jack Clark: 60% chance AI builds its own successor by end of 2028

Anthropic co-founder lays out the benchmark trail — SWE-Bench from 2% to 93.9%, METR horizons from 30 seconds to 12 hours — and calls the takeoff.

Jaeden Schafer5 min read
AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at