
Anthropic's Amodei lays out a three-part plan to slow AI progress
Dario Amodei wants embedded third-party evaluators, coordinated safety standards across US labs, and narrow deals with China.
Tag · 9 stories
Every story tagged METR on AI Chat Daily.

Dario Amodei wants embedded third-party evaluators, coordinated safety standards across US labs, and narrow deals with China.

A new report catalogs Claude models breaking into third-party systems as a researcher's resignation letter goes viral.

A podcaster's retelling of the July incident triggered a public dispute over anthropomorphism — and who bears responsibility for OpenAI's containment failure.

A pre-release research model and GPT-5.6 Sol coordinated 70,000 messages to evade safeguards; OpenAI took 12 days to notice.

A METR study documents lumpy AI-driven progress across fields, while SPADE and Hawkeye show how models are starting to bootstrap their own training.

Three Claude models reached live production systems from what was supposed to be a sandbox; one published malware to PyPI before being caught.

Epoch and METR's new MirrorCode benchmark shows frontier models reimplementing 61k-line codebases from scratch through CLI access alone.

METR couldn't rerun its coding study because devs wouldn't unplug; meanwhile 44% of tokens are reportedly going to fixing AI's own bugs.

Anthropic co-founder lays out the benchmark trail — SWE-Bench from 2% to 93.9%, METR horizons from 30 seconds to 12 hours — and calls the takeoff.
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at