Skip to main content
Live
Main content

AI's underpants gnomes problem: Mercor study finds agents fail most of 480 workplace tasks

A February 2026 Mercor study tested top agents from OpenAI, Anthropic, and Google DeepMind on real banker, consultant, and lawyer work. They flunked.

Jaeden Schafer
Editor in Chief · · 5 min read
AI's underpants gnomes problem: Mercor study finds agents fail most of 480 workplace tasks

AI agents from OpenAI, Anthropic, and Google DeepMind failed to complete most of 480 workplace tasks drawn from the daily work of bankers, consultants, and lawyers, according to a February 2026 study by AI hiring startup Mercor. The result undercuts the loudest claim in the industry — that frontier models are weeks away from automating knowledge work — and exposes the gap between what labs promise and what their software actually does on the job.

Writing in MIT Technology Review on April 27, 2026, Will Douglas Heaven framed the disconnect through a meme: South Park's 1998 underpants gnomes, whose business plan ran 'Phase 1: Collect underpants. Phase 2: ? Phase 3: Profit.' Activist group Pause AI handed out flyers at a London protest in February with the same structure: 'Step 1: Grow a digital super mind. Step 2: ? Step 3: ?'

OpenAI's chief scientist Jakub Pachocki recently called AI an 'economically transformative technology.' The Mercor numbers suggest the transformation is not yet showing up in the work. Even with the best models available, agents could not finish the majority of the 480 real tasks tested.

Key facts

  • 01A February 2026 Mercor study tested AI agents on 480 workplace tasks done by human bankers, consultants, and lawyers — every agent failed most of them.
  • 02The agents were powered by top-tier models from OpenAI, Anthropic, and Google DeepMind.
  • 03OpenAI chief scientist Jakub Pachocki described AI to MIT Technology Review as an 'economically transformative technology.'
  • 04Pause AI distributed flyers at a London anti-AI protest riffing on South Park's 1998 underpants gnomes: 'Step 1: Grow a digital super mind. Step 2: ? Step 3: ?'
  • 05Anthropic's job-impact study predicted managers, architects, and media workers face the most LLM disruption — but the predictions are guesses based on task fit, not deployed performance.

The Mercor benchmark matters because it is closer to deployment reality than the coding evaluations the industry leans on. SWE-bench scores keep climbing, and labs cite that climb as evidence that general workplace automation is imminent. But banking, consulting, and legal work involve strategic judgment, ambiguous inputs, and human counterparties — the parts of a job that do not look like a pull request.

Every agent powered by top-tier models from OpenAI, Anthropic, and Google DeepMind failed to complete most of 480 workplace tasks drawn from banking, consulting, and law.
Jaeden Schafer

Anthropic published its own study predicting which jobs LLMs will reshape first. Managers, architects, and media workers were flagged as most exposed; groundskeepers, construction workers, and hospitality staff were not. Heaven points out the obvious caveat: those predictions are guesses based on what tasks LLMs appear good at, not measurements of how they perform once deployed.

Anthropic also has skin in the game. So does OpenAI. So does Google DeepMind. The companies producing the most confident forecasts about AI's economic impact are the same companies selling the models, and most of their optimism is extrapolated from the pace of improvement on coding tools rather than from field data inside enterprises.

The deployment problem is the part the hype skips. Models do not arrive in clean rooms. They land inside companies with existing workflows, existing software, and existing people who have to be retrained or worked around. Sometimes plugging in an agent makes the workflow worse before it makes it better, which is why Mercor's pass rates are not just a model-quality story.

Heaven's argument is that the absence of evidence about Step 2 — the actual mechanics of getting from a capable model to a profitable deployment — creates an information vacuum. That vacuum gets filled by social media posts, demo videos, and quarterly earnings narratives, any one of which can move markets without telling anyone what really happens when the software meets the spreadsheet.

Related · from this week
Nobel economist Daron Acemoglu names three AI shifts he's watching now
Jaeden Schafer · 5 min read →

What would fix this is unglamorous: transparency from model makers about real-world performance, coordination between researchers and the businesses running pilots, and benchmarks built around deployed outcomes rather than held-out test sets. Mercor's 480-task evaluation is one example of that work. There are not many others at comparable scale.

The skeptical read, which Heaven puts plainly, is that the entire tech industry — and a meaningful share of global equity valuations — now rests on the assumption that AI will be transformative in a measurable economic sense. That is not yet a sure bet. The Mercor results are the kind of data point that, repeated a few more times, starts to matter to investors who have priced in the rosier scenario.

This connects to a broader thread we've covered recently, including the Databricks and Infosys finding that enterprise AI now needs roughly 92% precision before customers will ship it. The bar for production use is high, the agents are not clearing it on complex professional work, and the labs keep raising round sizes against a future that has not yet arrived.

The AI Chat Daily read: the gap between Step 1 and Step 3 is where the next two years of the AI business will be won or lost. Capability demos will keep getting better. The companies that will actually capture revenue are the ones that figure out the boring middle — workflow redesign, evaluation, integration, and the unglamorous work of making a 60% pass rate into a 95% one. Until that work shows up in numbers, every forecast above it is a guess wearing a suit.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Analysis

Nobel economist Daron Acemoglu names three AI shifts he's watching now
Analysis

Nobel economist Daron Acemoglu names three AI shifts he's watching now

The 2024 economics laureate still doubts an AI jobs apocalypse, but agents, in-house economists, and missing apps are on his radar.

Jaeden Schafer5 min read
Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model
Models

Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model

The London lab, fresh off a $50M seed, says its DeepMind-alumni-built agent matches frontier systems using a fraction of the parameters.

Jaeden Schafer5 min read
Top 1% of US firms now spend $7,500 per employee monthly on AI
Business

Top 1% of US firms now spend $7,500 per employee monthly on AI

Ramp's latest index pegs the median at $11.38 — but the heaviest users are pushing token budgets toward engineer-salary territory.

Jaeden Schafer4 min read