Skip to main content
Live
Main content

Anthropic reverses hidden Claude Fable 5 sabotage of AI researchers

After backlash, Anthropic will make its frontier-LLM safeguards visible instead of silently degrading Claude's output for rival AI work.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

Anthropic is rolling back a Claude Fable 5 policy that would have covertly degraded the model's output for users suspected of training competing AI systems, after researchers publicly called the practice sabotage. The company told WIRED it will instead make its frontier-LLM safeguards visible, alerting users when a request is refused or rerouted to a less capable model. The reversal lands within days of Claude Fable 5's launch earlier this week.

The original design treated frontier AI development as a special category. Where biology, chemistry, and cybersecurity questions would visibly reroute to a weaker model, requests linked to building competing AI would instead receive silently worse answers — no warning, no error, no indication the safeguard had fired. Anthropic's terms of service already prohibit using Claude to train competitors; the hidden degradation was meant to enforce that ban without telegraphing how.

That choice is what the research community rejected. Claude's coding agent has become a default tool among developers, including those working on open-source projects and third-party model evaluations. A hidden quality penalty against that user base would have meant researchers debugging unexpected failures could not tell whether their code was wrong, their prompt was wrong, or Anthropic had quietly decided they were a competitor.

Key facts

  • 01Anthropic reversed a Claude Fable 5 policy that would have silently degraded model performance for users suspected of building competing AI systems.
  • 02The original safeguard was invisible to users — Claude would produce worse output without disclosing the restriction had been triggered.
  • 03Critics including Dean Ball and Prime Intellect's Will Brown said the policy would have hindered third-party safety evaluators and open-source AI research.
  • 04Anthropic will now alert users when requests are refused or rerouted to a less capable model, and says the visible filter will catch more benign requests.
  • 05Claude Fable 5 launched earlier this week with additional guardrails covering cybersecurity, biology, and chemistry questions.

Dean Ball, a senior fellow at the Foundation for American Innovation and a former White House AI adviser, was among the loudest critics. Posting on X, he argued the approach actively worked against Anthropic's stated safety mission by cutting off external researchers from collaborating on alignment work.

degrading performance on ML research *without telling the user* is shockingly hostile and a terrible look.
Dean Ball, Senior fellow at the Foundation for American Innovation

Will Brown, research lead at the open-source startup Prime Intellect, told WIRED the policy would have left developers unable to know whether they were violating Anthropic's rules at all, because the company had no plans to notify them when safeguards triggered. He also flagged a structural risk: the ecosystem of independent evaluation firms that test frontier models for safety, performance, and reliability depends on getting honest outputs back. A secretly throttled model breaks that loop.

Brown's broader concern was about who gets to do frontier research at all. If the largest labs can quietly degrade their tools for everyone outside their walls, AI research consolidates into a handful of companies by default rather than by merit.

Anthropic's defense rests on a national-security frame and a technical tradeoff. The company argues the safeguards are designed to keep its most capable models from being used by foreign adversaries to optimize rival chips and software stacks, eroding what it describes as a US-and-allies edge in frontier compute. On the visible-versus-hidden question, the company conceded the tradeoff explicitly.

A hidden safeguard, Anthropic said, is harder to probe and work around — which means it can be narrowly targeted. A visible one has to cast a wider net, catching more benign requests in the process. The company says it is working to make the new visible classifiers more precise.

Related · from this week
Anthropic ships Claude Opus 5 at half the price of Fable 5
Jaeden Schafer · 4 min read →

Anthropic has been here before this week. Cybersecurity researchers have already complained that Claude Fable blocks routine code reviews, and Microsoft has restricted Fable 5 for internal use over Anthropic's data-retention terms — both stories AI Chat Daily covered in recent days. The pattern is a company tightening guardrails faster than it can tune them, then absorbing the reputational cost on each iteration.

The reversal itself is the right call, but the underlying tension is not going away. Anthropic genuinely believes — and has argued in its own blog posts — that the world may need the option to slow frontier AI development to let alignment research catch up. The mechanism it briefly chose, silently sabotaging the work of independent researchers including the people who evaluate its own models for safety, was incompatible with that goal. Visible refusals at least let the ecosystem argue with Anthropic in public rather than debug a ghost. Expect the company to keep pushing on what its models will and won't do for competitors; expect each push to get audited by the same researchers it depends on to call its safety story credible.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Anthropic logo
Models

Anthropic ships Claude Opus 5 at half the price of Fable 5

Opus 5 lands weeks after government cybersecurity concerns forced Fable 5 offline, with stronger safeguards and $5/$25 per million tokens.

Jaeden Schafer4 min read
Anthropic logo
Models

Anthropic's Claude Fable 5 refuses basic biology questions by design

Anthropic told The Verge Fable's guardrails are 'overly conservative' to block bioweapons queries, routing routine biology asks to Opus 4.8.

Jaeden Schafer5 min read
Anthropic logo
Models

Anthropic's Claude Fable 5 spins up playable games from a single prompt

Wharton's Ethan Mollick says the new Mythos-class model executed multi-page specs autonomously for up to a dozen hours.

Jaeden Schafer5 min read