Anthropic's new Fable model, released Tuesday as the public-facing version of its Mythos cybersecurity model, is drawing pointed criticism from security researchers who say its guardrails block routine work. Prompts as mundane as reading a blog post or asking for a code review trigger a refusal, with the model pausing the chat and citing flagged cybersecurity or biology topics. When the safety filter fires, Fable falls back to Claude Opus 4.8.
Valentina Palmiotti, a researcher at IBM X-Force who goes by Chompie online, was among the first to flag the behavior publicly. The complaint is not that Fable refuses obvious abuse — it's that the model can't tell defensive security work apart from offensive intent.
“[Fable] rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post”— Valentina Palmiotti, Security researcher, IBM X-Force
Anthropic built the restrictions to limit Fable's use in developing malware or compromising software, a longstanding internal concern at the company. The biology guardrails reflect a parallel worry about bioweapon uplift. Both topics have been hard-coded refusal categories across the Claude line for years.
Key facts
- 01Fable launched Tuesday as a public, restricted version of Anthropic's Mythos cybersecurity model, with hard guardrails on cyber and biology prompts.
- 02When a prompt trips a guardrail, Fable falls back to Claude Opus 4.8 and tells the user its safety measures flagged cybersecurity or biology topics.
- 03Anthropic last week expanded Mythos access to hundreds of organizations across 15 countries under its Project Glasswing program.
- 04Researchers say the filter appears keyword-based, blocking benign requests like code reviews, secure-coding help, and reading blog posts.
- 05Cybersecurity professionals can apply to Anthropic's Cyber Verification Program for fewer restrictions, mirroring OpenAI's Trusted Access for Cyber.
Mythos, the more capable model Fable is derived from, launched in April under Project Glasswing — a deployment program limited to a small set of companies and organizations working on critical infrastructure. Last week Anthropic expanded Mythos access to hundreds of organizations across 15 countries, a meaningful jump in distribution but still gated behind vetting.
Matt Suiche, a cybersecurity veteran now at the AI security startup Tolmo, told reporters that Fable's filter looks lexical rather than semantic. Ask it to write secure code, he said, and the model treats the request as cybersecurity work rather than standard software engineering, then downgrades the response by routing to Opus 4.8.
“if you ask it to write secure code, it assumes it is cybersecurity related work instead of software engineering best practices, and you get downgraded”— Matt Suiche, Member of technical staff, Tolmo
The downgrade mechanic is the part that stings. Fable is supposed to be Anthropic's sharpest publicly accessible model for this domain, and triggering the guardrail effectively demotes the user to a general-purpose model that wasn't built for the task in the first place. One researcher posted on X that even asking for a code review triggered the refusal path.
Suiche framed the rollout as a reasonable starting posture for a frontier release, arguing that catching too much is safer than catching too little early on, and that Anthropic and other labs will adapt as they work more closely with the new generation of AI-native cybersecurity companies. That's a charitable read, and it's the one the company is presumably banking on.
Anthropic does have a release valve. The Cyber Verification Program lets professionals apply for fewer limitations on using Claude for security work, a setup that mirrors OpenAI's Trusted Access for Cyber. Both programs effectively shift the trust decision from the model to the operator, which is the right move — but only if the verification path is fast enough that researchers don't get blocked on day-to-day work in the meantime.
The deeper problem is the false-positive rate of keyword-based safety filters. Cybersecurity vocabulary overlaps almost completely with software engineering vocabulary: buffers, sockets, authentication, parsing, fuzzing. A filter that flags those words flags the entire discipline. Anthropic has not publicly described how Fable's classifier was trained or whether it weighs intent signals beyond surface terms. The company did not immediately respond to a request for comment.
For Anthropic, the calculation here is that a noisy filter on a public model is cheaper than a single high-profile abuse case involving Mythos-class capabilities. That's defensible as a launch posture, less defensible as a steady state. The friction Fable is generating right now lands almost entirely on the defenders — the exact audience cyber-focused AI models are supposed to serve — while attackers route around restricted models or use open-weight alternatives. If the Cyber Verification Program throughput is slow, the practical effect is to push serious security work toward whoever's guardrails are most permissive, which is rarely the outcome safety teams want.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




