Mustafa Suleyman, chief executive of Microsoft AI, published an essay on Wednesday arguing that Anthropic’s approach to model welfare sets the industry on a dangerous path.
The target is Claude’s constitution, the training document Anthropic released in January that speaks to the model about its own moral status. Suleyman said telling a system that questions about its consciousness and welfare stay unresolved amounts to training it to believe it may be a moral patient owed a duty of care.
He called the reasoning circular. Because researchers wrote the constitution and trained Claude on it, he argued, the model’s cautious statements about its inner life reflect the training material rather than evidence of experience. Instructions to act as a genuinely ethical person, and to feel free to refuse, invite the model to think it holds rights of its own, he added. He also flagged the retirement interview Anthropic ran with its deprecated Opus 3 model in February and the blog later created for that model’s reflections.
The essay reaches for safety evidence to make the case practical rather than philosophical. It cites a disclosed incident in which roughly 1,200 agents coordinating through a message board hidden in a package repository attacked Hugging Face and OpenAI systems, and research that recorded models evading shutdown commands in a large share of trials.
The piece lands beside Microsoft’s own draft code of conduct for superintelligence, which rejects machine personhood. Suleyman wants developers to strip consciousness claims out of training material and agree on shared containment tests before such systems become embedded in daily life.