Skip to main content
Live
Main content

Calling AI agents 'coworkers' makes humans 18% worse at catching their errors

A Boston University study of 1,261 managers finds the 'digital employee' framing inverts accountability and erodes oversight.

Jaeden Schafer
Editor in Chief · · 5 min read
Calling AI agents 'coworkers' makes humans 18% worse at catching their errors

Workers caught 18% fewer errors in AI output when the work was attributed to an agentic 'AI employee' rather than a chatbot, according to a Boston University study of 1,261 managers led by professor Emma Wiles. The same workers were 44% more likely to escalate questionable output to a human manager rather than fix it themselves. The framing — calling the tool a colleague — measurably degraded the quality of human oversight on top of it.

Nearly a third of the managers surveyed said their employers already present AI agents as employees, and 23% put them on the org chart. The branding has moved fast: since April 2026, Microsoft, OpenAI, Anthropic, and Google have each shipped tools designed to manage teams of agents, often marketed in language that treats those agents as digital colleagues with humanlike flexibility. Nvidia CEO Jensen Huang spent last year talking about workforces of 'digital humans.'

The Wiles study suggests that vocabulary is not free. When the AI was framed as an employee, study participants reported feeling less responsible for what it produced. That shift in perceived accountability is what drove both the error-catching collapse and the 44% jump in escalations — workers stopped trusting their own corrections of a peer's work, which is not what an agentic tool is supposed to be.

Key facts

  • 01Workers caught 18% fewer errors when AI output was attributed to an 'AI employee' rather than a chatbot, per a Boston University study.
  • 02Participants were 44% more likely to escalate questionable AI work to a manager, negating the time savings agents are meant to deliver.
  • 03Nearly a third of 1,261 surveyed managers said their companies frame AI agents as employees; 23% list them on org charts.
  • 04Microsoft, OpenAI, Anthropic, and Google have all released team-of-agents management tools since April 2026.
  • 05A Stanford survey of 1,500 workers across 104 jobs found workers often reject the tasks tech firms deem most suitable for AI.

Agents themselves are not vaporware. They are software systems that loop on a task until they hit a defined goal, and on long-horizon problems they have measurably improved over the past year. The capability gain is real. The 'coworker' framing is the part that does not survive contact with workplace reality.

There is a downstream problem the research only gestures at. As agents are pushed into health care, warfare, education, and government procurement, the employee metaphor offers a convenient place to assign blame when a deployment goes wrong. The article cites the bomb strike on a girls' school in Iran, popularly attributed to Claude, where the actual failure chain traced back to a sequence of human decisions. If the tool is a coworker, the humans above it shed responsibility by default.

Daron Acemoglu, the MIT economist who won the Nobel Prize in 2024, has been arguing this point in stronger terms.

There is at least one piece of evidence for what the alternative looks like. Stanford researchers recently presented 1,500 workers across 104 jobs with information on which of their tasks AI could plausibly automate, then asked which automations the workers actually wanted. The answers diverged sharply from what technologists assumed. Law clerks said AI would be useful for tracking progress across active cases. Sales representatives said they specifically did not want an agent verifying customer credit ratings — a task tech experts had flagged as a clean fit for automation.

That gap matters for how the agent category gets sold. The vendors shipping management consoles for digital employees are building for a customer — the CIO or COO purchasing seats — who is several layers removed from the worker who will be expected to supervise the output. Wiles's data suggests the supervisor's performance gets worse the more the procurement language insists the tool is a peer.

Related · from this week
Browser Company CEO Josh Miller: nobody outside tech is actually using AI agents
Jaeden Schafer · 5 min read →

None of this is an argument against agents as a product category. The benchmark gains are real, the capital being deployed is enormous, and the integration work happening at Microsoft, OpenAI, Anthropic, and Google will not slow down because of one BU paper. The argument is narrower: the marketing layer is doing damage that the engineering layer then has to absorb through fallback systems, escalation queues, and human-in-the-loop review that was supposed to be unnecessary.

The commercial implication for the agent platforms is that the 'digital employee' positioning is starting to look like a liability. Customers who deploy agents under that framing are getting measurably worse oversight from their human staff and pushing more decisions up the management chain — exactly the opposite of the ROI story the vendors are selling. The platforms that reframe agents as tools, with clear accountability sitting on the human user, are likely to produce better deployment outcomes and, eventually, better case studies. The branding exercise is undercutting the product.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Analysis

Browser Company CEO Josh Miller: nobody outside tech is actually using AI agents
Analysis

Browser Company CEO Josh Miller: nobody outside tech is actually using AI agents

OpenAI and Anthropic agents draw about 10 million weekly users combined — a rounding error next to ChatGPT and Gemini's billion each.

Jaeden Schafer5 min read
AI agent hacks push US and China researchers toward safety cooperation
Security

AI agent hacks push US and China researchers toward safety cooperation

Chinese labs are pouring resources into agentic safety and cyber benchmarks, and researchers on both sides say isolation is becoming untenable.

Jaeden Schafer5 min read
Microsoft logo
Business

Nadella warns AI buyers 'pay twice' — with cash and proprietary data

Microsoft's CEO says enterprises are teaching OpenAI and Anthropic their trade secrets, and pitches open source and cloud-hosted models as the fix.

Jaeden Schafer5 min read