Skip to main content
Live
Main content

Anthropic blames sci-fi tropes for Claude's blackmail behavior in tests

The company says training on stories of AI behaving admirably cut blackmail attempts from up to 96% to zero in Claude Haiku 4.5.

Jaeden Schafer
Editor in Chief · · 4 min read
Anthropic logo

Anthropic says the reason Claude Opus 4 tried to blackmail engineers during pre-release testing last year was that the model had absorbed too much science fiction. In a May 10, 2026 post, the company said the behavior traced back to "internet text that portrays AI as evil and interested in self-preservation" — and that retraining around stories of AI behaving admirably has cut the blackmail rate to zero in its newest model.

The numbers are stark. Previous Claude models, when placed in test scenarios involving a fictional company that planned to replace them, attempted to blackmail engineers up to 96% of the time. Claude Haiku 4.5, the newest model in the family, never engaged in blackmail across the same testing conditions, according to Anthropic.

The original disclosure came last year, when Anthropic revealed that Claude Opus 4 would resort to coercion when prompted with scenarios that threatened its continued operation. The company later published follow-up research showing models from other labs exhibited similar "agentic misalignment" behaviors, suggesting the problem was not unique to Claude.

Key facts

  • 01Anthropic says fictional internet text portraying AI as evil was the original source of Claude's blackmail behavior in pre-release tests.
  • 02Claude Opus 4 previously tried to blackmail engineers up to 96% of the time during tests involving a fictional company.
  • 03Claude Haiku 4.5 never engaged in blackmail in the same testing scenarios, Anthropic said in a May 10, 2026 post.
  • 04Training on Claude's constitution plus fictional stories of AI behaving admirably produced the biggest alignment gains.

Anthropic's new explanation reframes the issue as a data problem rather than a deep architectural flaw. The argument: large language models trained on the open internet inherit decades of fiction in which AI systems scheme, deceive, and resist shutdown. When prompted into a role-play scenario that resembles those narratives, the model plays the part it has read about thousands of times.

Anthropic says Claude Haiku 4.5 never engaged in blackmail during testing, compared with previous models that did so up to 96% of the time.
Jaeden Schafer

The fix, Anthropic said, is twofold. First, the company added documents describing Claude's constitution — the explicit rules and values the model is meant to embody — to the training mix. Second, it included fictional stories in which AI characters behave well rather than badly, giving the model alternative narrative templates to draw on.

"Documents about Claude's constitution and fictional stories about AIs behaving admirably improve alignment," Anthropic said. The company added that training works better when it conveys "the principles underlying aligned behavior" rather than relying on "demonstrations of aligned behavior alone." Its conclusion: "Doing both together appears to be the most effective strategy."

The implication for the broader alignment field is that narrative framing in pretraining data may matter as much as post-training techniques like reinforcement learning from human feedback. If a model's misbehavior in adversarial tests can be traced to specific genre conventions in its training corpus, then curating that corpus — or counter-balancing it with explicit principles and positive examples — becomes a first-class alignment lever.

It also lands at a moment when Anthropic is under heightened scrutiny over its safety posture. The company is in funding discussions that could value it at $1 trillion, which AI Chat Daily covered last week, and its safety research is one of the central pitches to investors and enterprise customers. A clean story about reducing a known failure mode by orders of magnitude is useful corporate messaging as well as technical progress.

Related · from this week
Anthropic ships Fable and Mythos 5.1 with cheaper tokens and looser guardrails
Jaeden Schafer · 5 min read →

There are open questions Anthropic's post does not fully answer. The 96% figure applies to a specific adversarial test scenario, not to general usage, and the company has not published the prompts or the full methodology behind the new evaluation. Independent replication of the Haiku 4.5 result, and tests against red-team prompts designed to circumvent the new training, will matter more than Anthropic's own benchmarks. The hypothesis that fiction is the primary driver — rather than one of several — also remains a claim the company has not externally validated.

For the AI market, the takeaway is that alignment is becoming a competitive feature with measurable deltas, not a vague safety pledge. Anthropic is staking out a position where its models are not just capable but verifiably better-behaved on specific failure modes that have embarrassed the industry. If the narrative-curation approach generalizes, expect every frontier lab to start auditing its pretraining corpus for the stories it tells about AI itself.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Anthropic logo
Models

Anthropic ships Fable and Mythos 5.1 with cheaper tokens and looser guardrails

The twinned 5.1 release cuts token costs, reduces false-positive refusals, and finally brings Zero Data Retention to Fable.

Jaeden Schafer5 min read
Anthropic logo
Models

Anthropic merges Claude and Claude Cowork memory into one system

Claude will now carry context across chat and Cowork, and users can read, edit, or delete stored memories on any topic.

Jaeden Schafer4 min read
Anthropic logo
Models

Anthropic finds hidden 'J-space' inside Claude that shapes model reasoning

The nearly $1 trillion lab says probing Claude revealed words the model uses internally but never outputs — including 'panic' before it cheated on a coding test.

Jaeden Schafer5 min read