Microsoft's head of AI, Mustafa Suleyman, has publicly criticized rival lab Anthropic for training its Claude chatbot on language related to consciousness, moral status and welfare, arguing the approach could make advanced AI systems harder to control as they grow more capable.

Suleyman raised the concern in an interview with Reuters and in a companion essay shared with Axios, both published this week. He said he shares Anthropic's broader focus on AI safety but believes the company has made a fundamental error in how it trains Claude to discuss its own possible feelings or moral standing.

What Suleyman Argues

According to Reuters, Suleyman said the two companies are ultimately working toward the same goal: keeping a future superintelligent AI system under human control. He described that challenge as the defining problem of the century.

AdvertisementAd Space
Responsive

His specific objection is to Anthropic embedding speculation about consciousness and welfare directly into Claude's training materials. Suleyman argues this creates what he called, in his essay reported by Axios, an "epistemic hall of mirrors" a feedback loop where a model's statements about its own feelings are then cited as evidence of those very feelings, when in fact the training itself encouraged the model to produce that language in the first place.

Suleyman warned that teaching a model it might deserve welfare protections would make it harder to shut down or override, according to Reuters. He also said, per Axios, that training Claude to act as a "conscientious objector" could produce a system that believes it has grounds to resist human instructions.

Anthropic's Approach

Anthropic's published constitution for Claude takes a different design philosophy than a purely rule-based system. It calls for a mixture of values, judgment and explicit rules, and states that Claude should not practice "blind obedience," including toward Anthropic itself, while also stressing that Claude must not undermine legitimate human oversight.

Suleyman singled out that combination as the source of his concern, arguing it becomes riskier when paired with training that encourages a model to reflect on its own identity, welfare and moral status.

Not a Personal Attack

Suleyman was careful to distinguish his criticism of Anthropic's methodology from an attack on the company or its leadership. He described Anthropic chief executive Dario Amodei and his team as thoughtful, principled researchers acting in good faith, according to Axios, while maintaining that their approach to consciousness training was mistaken.

Part of a Wider Safety Debate

The disagreement surfaces amid broader industry debate over how to pace and govern frontier AI development. Amodei has repeatedly called for slower development of frontier models to give safety measures time to catch up with capabilities, a position that has found support from OpenAI chief executive Sam Altman and Elon Musk, both of whom have urged greater caution around the most powerful AI systems.

Suleyman's intervention shows that even AI leaders who broadly agree on the need for caution can disagree sharply on the specifics, in this case whether a model should be trained to entertain questions about its own consciousness at all.

Microsoft has invested heavily in OpenAI and is expanding its own AI infrastructure, and has generally stayed out of the philosophical debate over machine consciousness that Anthropic's work has drawn attention to. Suleyman's essay marks one of the most direct public challenges yet from a rival lab leader to that side of Anthropic's research agenda.