
AnalysisLead story
AI's underpants gnomes problem: Mercor study finds agents fail most of 480 workplace tasks
A February 2026 Mercor study tested top agents from OpenAI, Anthropic, and Google DeepMind on real banker, consultant, and lawyer work. They flunked.
Jaeden SchaferEditor in Chief5 min read