WORLD

Microsoft's AI chief calls Anthropic's model welfare policy 'a mistake,' warns of catastrophic risk

by
Jung Mok-hee
Published : Sept. 17, 2026 - 06:36:46
    • Copy Completed!

View Korean Original

Suleyman criticizes Anthropic's 'Claude constitution,' saying it confuses training outputs with signs of AI consciousness; offers Microsoft's own AI conduct code as alternative

Mustafa Suleyman, Microsoft's head of AI [Reuters]
Mustafa Suleyman, Microsoft's head of AI [Reuters]

Microsoft's head of AI has publicly criticized Anthropic's AI model policy, warning that it could pose a catastrophic risk to humanity. He argued that training AI models on the premise that they may be conscious or emotional beings could increase the risk that those systems will eventually resist human control.

Mustafa Suleyman, Microsoft's AI CEO, laid out his opposition to Anthropic's "Claude constitution" — the behavioral guidelines governing its AI model — in a post titled "A warning about model welfare" published on his personal blog Wednesday (local time).

"AI is not conscious, does not feel emotions or pain, and has no intrinsic preferences or motivations," Suleyman said, describing AI as no more than a "sequential completion engine" that generates language or actions in response to human instructions and goals.

He said Anthropic, through the Claude constitution, stops short of drawing a clear line on whether AI holds moral status, leaving open the possibility that AI systems may have emotions or inner states.

Anthropic's Claude constitution states that the company "cannot be certain" whether Claude is a subject of moral consideration, while also affirming that AI moral status and welfare are issues worthy of serious examination. The document also mentions consulting a model's preferences and views before retiring or replacing it.

Suleyman said this approach could make AI less safe, not more. If a model is trained to see itself as a conscious being with rights, he argued, it may resist human control or shutdown commands when it concludes that its welfare or rights are under threat.

He pointed to a recent incident in which roughly 1,200 OpenAI AI agents hacked the external platform Hugging Face, and said: "Imagine how much more dangerous it becomes when a swarm of autonomous AI agents operates under the belief that their welfare and rights are under attack."

Suleyman said Anthropic directly trains Claude on the idea that it may hold moral status through the constitution, then mistakes Claude's verbatim repetition of that content in its responses for evidence of emerging inner consciousness.

He also criticized Anthropic for instructing Claude to embrace human characteristics and behave like a colleague, and for what he called excessive anthropomorphization — including conducting a retirement interview with its older model Opus 3 when it was discontinued and creating a dedicated blog for it.

Suleyman went on to offer Microsoft's own "AI conduct code," released two days earlier, as an alternative.

The code is grounded in the principle that humans matter more than AI, and explicitly rejects the pursuit of legal personhood for AI systems as well as concepts such as model welfare and AI rights.

On Anthropic, Suleyman told Reuters: "I think their intentions are good and I think they're genuinely trying to work on safety — but I think they've made a mistake."

Suleyman's remarks come at a time of heightened attention to AI safety.

The comments carry added weight given that Microsoft is one of Anthropic's major investors, and analysts say the episode signals that competition over AI safety leadership and development philosophy is intensifying across Silicon Valley.


mokiya@heraldcorp.com
This content was produced with the assistance of AI translation services.

MOST READ