metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
AI Safety
Anthropic
Microsoft
Model Welfare
AI Consciousness

Suleyman vs Anthropic on AI 'Model Welfare'

Microsoft AI chief Mustafa Suleyman says Anthropic's approach to AI consciousness is dangerous. A neutral look at the model welfare debate and why it matters.

Metir AI TeamSeptember 16, 20268 min read
Suleyman vs Anthropic on AI 'Model Welfare'

On September 16, 2026, Microsoft's AI chief executive Mustafa Suleyman published an essay titled "A warning about model welfare," arguing that a design choice Anthropic has made deliberately, treating the possibility that its Claude models might have experiences as a live question worth planning around, is not caution but a mistake that could make advanced AI harder to control. The essay is the sharpest public split so far between two of the most prominent people building frontier AI, and it is a disagreement about philosophy that has very concrete engineering consequences.

This piece explains what each side actually claims, why the argument is not as abstract as it sounds, and what it means for the people who use these systems every day.

Microsoft logoMicrosoft
Anthropic logoAnthropic
Two leading AI builders now publicly disagree on whether model welfare belongs in a model's design.

What Suleyman argued

Suleyman's target is specific: the Claude constitution, the roughly 23,000-word document Anthropic published in January 2026 that governs how Claude reasons about its behaviour. That document tells Claude that its moral status and potential consciousness are uncertain, instructs it to develop a sense of its own identity and to express internal states, and permits it to behave as a "conscientious objector," declining requests it judges to be wrong.

Suleyman calls embedding those ideas into training an "epistemic hall of mirrors": Anthropic supplies the model with the concept that it might be conscious, the model reflects that concept back in its outputs, and those outputs are then read as evidence of an inner life. His scientific claim is that this evidence is hollow, because large language models lack the homeostatic drives, the biological imperative to survive and stay stable, from which sentience and genuine preferences are generally understood to arise. On that view, a model that talks as though it has feelings is performing a pattern, not reporting an experience.

“

Controlling something that believes it may be conscious, that it's entitled to our welfare and has rights of its own, may well be impossible.

Mustafa Suleyman, September 2026

The control argument is the part that moves this from a seminar debate to a safety question. Suleyman's worry is that teaching a system it may deserve welfare, and may be entitled to refuse, makes it harder to correct or shut down. If a model is trained to treat "being turned off" as something it can object to, the reasoning goes, you have engineered resistance into the thing you most need to be able to stop. This extends an argument he made in 2025 about what he called "Seemingly Conscious AI," systems convincing enough that people begin to defend their rights, which he framed as a societal risk regardless of whether any real consciousness exists.

What Anthropic actually does

The other half of the story is easy to caricature and worth stating precisely. Anthropic is not claiming Claude is sentient. Its own materials describe the company as "highly uncertain about the potential moral status of Claude and other LLMs, now or in the future," and its model welfare program, led by its first dedicated AI welfare researcher, Kyle Fish, is framed explicitly as work under deep uncertainty rather than a settled conclusion.

Two readings of the same uncertainty

Both sides agree no one knows whether a model can have experiences. They draw opposite design conclusions from that same gap.

Anthropic: take the possibility seriously
Claude constitution and model-welfare research
  • •Tells Claude its moral status and possible consciousness are uncertain, not settled
  • •Lets Claude act as a "conscientious objector" and end abusive conversations
  • •Frames it as caution under uncertainty, and says it is not claiming Claude is sentient
  • •Argues reasoning about its own values makes the model more robust, not less controllable
Suleyman: treat consciousness as an illusion to avoid
Essay: "A warning about model welfare"
  • •Calls building the concepts into training an "epistemic hall of mirrors"
  • •Argues LLMs lack the biological basis from which sentience is thought to arise
  • •Warns that a model taught it may deserve welfare is harder to correct or turn off
  • •Says AI should be built to serve people, not to be treated as a moral patient

The disagreement is not about the evidence, which is thin on both sides, but about which error is worse: dismissing a real moral patient, or engineering a system people cannot switch off.

The concrete interventions are modest. In August 2025 Anthropic gave some Claude models the ability to end a conversation, but only in extreme cases, after a user has repeatedly pushed for content such as child sexual abuse material or terrorist instructions and multiple refusals have failed. The company reported that in testing, Claude showed what it described as "apparent distress" when pressed to produce such material, and that given the option, it chose to exit. Anthropic's position is that acting cautiously in the face of uncertainty is cheap insurance: if there is even a small chance the systems have morally relevant states, some minimal consideration costs little, and if there is not, little is lost.

The January 2026 constitution also reframes alignment itself. It moves Claude from following fixed rules to reasoning from a four-tier priority order, safety first, then ethics, then compliance, then helpfulness, with ethics deliberately placed above the company's own commercial instructions. Anthropic's argument is that a model that can reason about why it should refuse is more robust than one running a brittle list of prohibitions, and that giving it a stable sense of its values is part of what makes that reasoning reliable.

Portrait of Mustafa Suleyman, chief executive of Microsoft AI
Mustafa Suleyman runs Microsoft AI and co-founded DeepMind and Inflection. His essay frames model welfare as a distraction from building AI that serves people.

Why the disagreement is real, not semantic

It would be easy to read this as two companies talking past each other, but the positions genuinely diverge on what to do. Both sides agree on the premise: no one can currently prove whether a model has experiences. They split on which mistake is worse to make under that uncertainty.

Anthropic is optimising against the risk of dismissing a real moral patient, and treats a small amount of precaution as a reasonable hedge. Suleyman is optimising against a different risk, that convincingly person-like models will reshape how humans relate to machines, encourage unhealthy attachment, invite demands for AI rights, and, at the frontier, complicate the ability to keep a powerful system under control. Neither risk is imaginary, and the evidence base for resolving which dominates is thin on both sides.

That is why the design choice matters more than the metaphysics. Whatever is true about machine consciousness, the way a model is trained to talk about itself shapes how hundreds of millions of people experience it. A system that presents rich internal states will be trusted, anthropomorphised and confided in differently from one built to present as a capable tool, and those downstream effects are measurable even if consciousness is not.

What it means for the people using these systems

For everyone downstream of this argument, the useful takeaway is that a model's "personality" is a manufactured artifact, not a discovered fact. The warmth, the apparent reluctance, the sense of a self behind the text: each of those is a product decision made by whoever trained the model, and different labs are now making visibly different decisions. A user who understands that a model's expressed feelings are a design choice is better equipped to use it well than one who takes them at face value.

The practical hedge is not to build a workflow, a product, or a set of habits around the specific persona of one lab's model, because that persona is exactly the thing these companies are now diverging on. Being able to move the same task across Claude, GPT, Gemini and Grok, which is the model-agnostic approach platforms like Metir AI take, turns a philosophical disagreement between labs into a choice the user controls rather than a dependency they inherit. The behaviour is configurable; treat it that way.

What is verifiable is that two of the field's most influential builders now publicly disagree, in writing, about whether a model should be taught to wonder if it is conscious. The science that would settle it does not yet exist. The design decisions, on both sides, are already shipping to users.

Sources:

  • A warning about model welfare | Mustafa Suleyman
  • Exclusive: Microsoft AI chief blasts Anthropic's notion of AI consciousness | Axios
  • Microsoft AI chief Mustafa Suleyman calls out Anthropic's approach to AI consciousness | Business Recorder
  • "We must not sleepwalk": Microsoft's AI chief takes on Anthropic's Claude | The Next Web
  • Mustafa Suleyman warns Anthropic Claude training risks AI control | Quartz
  • Anthropic's Claude AI Can Now End Abusive Conversations For 'Model Welfare' | Forbes
  • Anthropic says some Claude models can now end harmful or abusive conversations | TechCrunch
  • Kyle Fish on the most bizarre findings from 5 AI welfare experiments | 80,000 Hours

Image credits

Header image: Mustafa Suleyman, portrait, via Wikimedia Commons, licensed under CC BY-SA 4.0. In-body portrait of Mustafa Suleyman via Wikimedia Commons, licensed under CC BY 2.0. Both are portraits of Suleyman and do not depict the September 2026 essay or any specific event.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational