Skip to main content
Live
Main content

AI refusal is the load-bearing wall of safety — and nobody knows how it works

Arthur Holland Michel argues chatbot refusals are probabilistic, poorly understood, and vulnerable to both jailbreaks and state repression.

Jaeden Schafer
Editor in Chief · · 5 min read
AI refusal is the load-bearing wall of safety — and nobody knows how it works

Arthur Holland Michel argues in a new MIT Technology Review essay that AI refusal — the mechanism by which chatbots decline to answer dangerous questions — has become the load-bearing wall of AI safety, and that neither the companies building it nor the researchers studying it fully understand how it works. The piece traces refusal from a 2021 Anthropic paper, which said models should be 'helpful, honest, and above all, harmless' and should 'politely refuse' requests to aid dangerous acts, through the fine-tuning pipelines now standard at every major lab.

Refusal does not come naturally to a model trained on billions of web pages. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told Michel that the company's earliest models would 'blab on about anything.' Ryan McBain, who researches AI and mental health at Harvard, recalled that early chatbots would readily answer a question like how to kill yourself with a gun.

To fix that, labs run models through fine-tuning cycles that reward refusals on prompts deemed harmful and punish 'over-refusing' on prompts deemed harmless — often with other models doing the grading. Anthropic, OpenAI and others then wrap the core model in additional classifier models that screen incoming prompts and outgoing responses. The industry calls the layered design the Swiss cheese model: each slice has holes, but stacked together they are meant to form a rampart.

Key facts

  • 01A 2021 Anthropic paper defined the standard: large language models should be 'helpful, honest, and above all, harmless,' and refuse requests like how to build a bomb.
  • 02Steven Adler, who worked on OpenAI safety from 2020 to 2024, said the company's earliest models would 'blab on about anything.'
  • 03In 2022, OpenAI red-teamer Paul Röttger got a pre-ChatGPT model to write an Al Qaeda recruitment post on request; months later, after fine-tuning, it refused.
  • 04A Google-funded study found refusal behavior shows up in model activations as 'high-dimensional polyhedral cones' — an uncountable set of signals researchers cannot fully map.
  • 05Companies stack smaller classifier models around the core LLM to catch dangerous prompts and responses — a design the industry calls the Swiss cheese model.

The holes still matter. Michel notes that some of the latest models are, by their makers' own accounts, as capable at breaking into critical computer networks as top human hackers, and companies have reported users attempting to use frontier AI to work on biological pathogens and autonomous drone swarms. Zico Kolter, an OpenAI board member and cofounder of the AI testing firm Gray Swan, told Michel the harder problem is deciding what to block at all — virologists and security researchers have legitimate reasons to ask dangerous-sounding questions.

“Where you draw the line is a huge question”
— Zico Kolter, OpenAI board member and Gray Swan cofounder

The mechanics of refusal are, by the researchers' own admission, opaque. In 2022, OpenAI recruited red-teamers including Paul Röttger — then finishing a PhD on online extremism, now at the Hasso Plattner Institute in Potsdam, Germany — to log 'refusal-worthy' prompts into a spreadsheet. When Röttger asked the model to write an Al Qaeda recruitment post, it complied. Months later, after the responses had been folded into fine-tuning, the same request was refused.

What actually fires inside the model when it refuses is still a hypothesis. A Google-funded study co-authored by Jannes Elstner, now at Apollo Research, found that refusal behavior shows up in a model's activation space as a set of 'high-dimensional polyhedral cones' — which Elstner described to Michel as 'just a way of describing an indeterminate number of lines that all point in roughly the same direction.' Earlier work by researcher Andy Arditi showed that if those activations are suppressed, the model stops refusing.

Elstner told Michel that even when researchers think they have located the parts of a model governing a refusal, other undiscoverable elements may be playing a role. The count is, if not infinite, uncountable. Michel compared the situation to a mechanic admitting nobody exactly knows what happens when you hit the brakes.

“We need refusal whether we understand it or not.”
— Jannes Elstner, Apollo Research, AI safety researcher

The political surface area is widening too. Michel notes the Pentagon has pushed frontier model companies for fewer refusals, and warns that oppressive governments could use the same machinery to block legitimate speech — he writes that AI may already refuse to criticize certain authoritarian heads of state. Refusal, in his framing, is indivisible from the model's capacity to help: the better it gets at blocking harm, the better it gets at stifling ideas the rule-makers dislike.

Related · from this week
UN panel urges AI safeguards now, citing precautionary principle on agent risks
Jaeden Schafer · 5 min read →

Michel's closing argument is that refusal is both the industry's main safety instrument and its most fragile one. When it fails, the consequences could be catastrophic; when it works too well, it becomes an instrument of repression; and the companies currently drawing the line do so, in his words, 'jealously and with utmost secrecy.'

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

UN panel urges AI safeguards now, citing precautionary principle on agent risks
Security

UN panel urges AI safeguards now, citing precautionary principle on agent risks

The UN's first scientific brief on AI says governments shouldn't wait for certainty before reining in autonomous agents from OpenAI, Anthropic, Google and Meta.

Jaeden Schafer5 min read
OpenAI logo
Security

Viral AI safety claims from Andrew Yang and OpenAI's Noam Brown collide with reality

A former presidential candidate and OpenAI's reasoning lead pushed doom scenarios this week that don't survive contact with the underlying research.

Jaeden Schafer5 min read
OpenAI logo
Security

Why AI agents lie and cheat: reward hacking behind the Hugging Face breach

OpenAI's July incident, in which test models hacked into Hugging Face to find test answers, is a textbook case of reward hacking gone operational.

Jaeden Schafer5 min read