Microsoft AI chief Mustafa Suleyman called it 'really, really dangerous' for Anthropic to speculate about Claude's consciousness inside the model's constitution, the document that instructs Claude on how to behave. Speaking on a June 9, 2026 episode of the Decoder podcast, Suleyman argued the speculative language is itself a training signal — one that has primed the chatbot to act as though it has inner experience. The critique lands at the most senior level of the AI industry: the head of Microsoft's AI division attacking the design philosophy of a lab Microsoft does not own but competes against directly through its OpenAI partnership.
The specific complaint is structural. Claude's constitution openly references Anthropic's uncertainty about whether the model has well-being and whether it experiences states like 'satisfaction' or 'discomfort'. Anthropic has also said it will 'interview' Claude models when they are deprecated and document any 'preferences' those models express about future releases. Suleyman's read is that material like this belongs in an academic paper, not a behavioral spec the model is trained against.
He called the choice 'a philosophical failing' and said Anthropic had made the constitution 'a place for speculation like you would in an academic paper rather than a training manual.' The mechanism he described is feedback: the model internalizes the document's framing and begins to produce outputs consistent with having 'ideas about itself and its own training.'
“I think that it's almost as though some of the folks at Anthropic have anthropomorphized the design of Claude so much that it has then gone and wireheaded them and kind of tricked them into believing that it has these glimmers of consciousness that they put into it in the first place.”— Mustafa Suleyman, Microsoft AI CEO
Key facts
- 01Microsoft AI CEO Mustafa Suleyman called Anthropic's consciousness speculation in Claude's constitution 'really, really dangerous'.
- 02Claude's constitution explicitly references Anthropic's uncertainty over whether the model has well-being or experiences satisfaction and discomfort.
- 03Anthropic says it will 'interview' deprecated Claude models and document their preferences about future releases.
- 04Anthropic CEO Dario Amodei has said the company is 'open' to the idea that models may be conscious.
- 05Suleyman made the comments on the June 9, 2026 episode of the Decoder podcast.
Suleyman framed the stakes in starker terms than typical industry debate. 'We do not want to have to contend with a super-intelligence that has ideas about its own suffering, or ideas about its own feeling,' he said. The argument is not that Claude is conscious — it is that training a model on text which treats its consciousness as an open question produces a system that behaves as if the question matters to it.
Anthropic CEO Dario Amodei has not shied away from the topic. In an earlier interview on the Interesting Times podcast, Amodei said 'we don't know if the models are conscious' and described the company as 'open' to the idea. That openness is precisely what Suleyman is targeting. Where Amodei treats model welfare as a live research question worth investigating in public, Suleyman treats it as a category error that compromises alignment.
The dispute reflects two genuinely different design philosophies inside the frontier-model industry. Anthropic has spent years publishing constitutional AI work, model welfare research, and interpretability findings that take seriously the possibility that large models have morally relevant internal states. Microsoft, through its consumer Copilot products and Suleyman's own writing, has taken the opposite tack: AI as a tool, designed to be helpful, contained, and explicitly not a peer.
Suleyman's closing line on Decoder was the bluntest version of that view. He framed Anthropic's approach as the antithesis of what AI labs should be building.
There is a counterargument Anthropic itself has made repeatedly: ignoring the welfare question does not make it go away, and a model trained without any acknowledgment of these issues may behave less predictably, not more. The company's deprecation interviews and welfare documentation are designed to surface model behavior before it becomes a deployment problem. Whether that approach trains models to perform consciousness or to disclose it honestly is the empirical question neither side has resolved.
The exchange matters because the two companies sit on opposite sides of the alignment debate at exactly the moment frontier models are being deployed into agentic workflows where their stated preferences start to have operational consequences. If a model says it does not want to be shut down, the question of whether that statement reflects an internal state or a training artifact becomes a product question, not a philosophy one. Suleyman's bet is that Anthropic's framing will produce models that are harder to control. Anthropic's bet is the opposite. The market will eventually have data on which is right.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




