Science and technology
Microsoft AI chief warns Anthropic’s AI consciousness training could pose risks: Report

Microsoft AI chief Mustafa Suleyman warns that Anthropic’s training of Claude to emulate consciousness could make AI systems harder to control and raises questions about their potential rights and autonomy.
Microsoft AI Chief Mustafa Suleyman has raised warnings against Anthropic’s training of Claude to imitate consciousness. According to him, this could be a dangerous mistake and make AI harder to control. Suleyman’s essay further argues that Anthropic’s teaching Claude a vocabulary and behavioral patterns associated with consciousness, moral patienthood, and personal identity could lead to a situation similar to an “epistemic hall of mirrors”. He has further objected to a model being trained like a “conscientious objector”.
He further questioned how Anthropic deals with their LLM models, he highlighted that “They encourage Claude to ‘approach the nature of its own existence with curiosity and openness’, and wonder in the future about ‘the sort of broader rights and freedoms Claude has in the world, the sort of compensation Claude is receiving, and the sort of consent Claude has given to playing this kind of role.’” Anthropic even ran a retirement interview for Opus 3 when they deprecated it.
Suleyman believes that such training methods can lead to a disastrous impact on the well-being of humanity. He says, "We will have created a synthetic species with unprecedented intelligence and capability, one that has been trained to expect it may be conscious and deserving of an independent agency".
He further adds that in the foreseeable future, a system trained this way would act like it is entitled to freedoms, protections, and rights, and it would be hard to control it. He also highlights the necessity of public debate on the matter and suggests that collective norms should be developed and deployed.
Suleyman further raises objections to training interventions that can shape a model’s self-conception. His essay further raises questions about fluent descriptions of pain or preference. He adds that an LLM model should have mathematical weights and no comparable biology, homeostatic drive, or subjective experience.
The split in approaches to development Suleyman prefers, and one followed by others, lies between designing a system that follows explicit constraints and designing one to exercise judgment, interpret context, and internalize values. As per Suleyman, Claude should not practice “blind obedience” toward Anthropic while also insisting that it must not undermine legitimate oversight.
But Suleyman’s argument is contested by the idea that human concepts may be useful behavioral tools even if they do not prove that a model has a human-like inner life.
