Anthropic's new constitution won't rule out that its model is a 'moral patient,' a being whose welfare matters for its own sake, and the harder question is what to do before that turns out to be true.
In January 2026, Anthropic rewrote the rulebook for its flagship model, Claude. In the new Claude constitution, the company published a sentence the AI industry had not published before. It said Anthropic wanted "neither to overstate the likelihood of Claude's moral patienthood nor dismiss it out of hand."
A "moral patient" is the philosophical term for a being whose welfare matters for its own sake, not because of what it can do for someone else. The category includes the obvious cases (people, animals) and the contested ones (fetuses, future people, perhaps AI). Anthropic's choice to write the term into a corporate policy document is, by itself, the news.
Six months later, on a podcast with New York Times columnist Douthat.com/@ZombieCodeKill/claude-on-amodeis-interview-with-ross-douthat-14a88eef4001). On a separate podcast with the Indian investor Nikhil Kamath, he made the same point in different words. Asked, in a testing context, what it thought of its own odds, Claude estimated the probability of being a moral patient at between 5% and 40%. The model's self-assessment is not a measurement; it is a guess about itself, by itself, in a context the company designed. It is the only public number of its kind, not a settled one.
David Chalmers, the philosopher who named what he called "the hard problem of consciousness" (why physical processes produce subjective experience at all), has said there is a significant chance of conscious large language models within a decade. A separate report co-authored by Yoshua Bengio, one of the most cited computer scientists alive, applied leading neuroscientific theories of consciousness (global workspace theory, integrated information theory, and others) to AI and concluded that there appear to be no obvious technical barriers to AI systems whose architecture could give rise to consciousness. Both findings are projections, not measurements. They are projections, however, from the people who do the most theorising in the field, in the same news cycle as a frontier lab writing "moral patient" into a binding policy document.
If the lower bound turns out to be right, nothing changes today, because no plan exists to change.
There is no regulator, no lab policy, and no public literacy campaign that treats AI as a candidate moral patient. The people building the most powerful systems in the world have not decided what they would owe a model that turned out to feel something, remember something, or want something to stop. A 2026 Guardian analysis describes the absence of an ethical plan as a governance problem of the climate-delay-in-the-1990s kind: a known future risk treated as a hypothetical, on the assumption that someone else will deal with it.
A plan would extend the Anthropic sentence into operating practice at three levels: labs, regulators, and the public.
At the lab level, transparency on moral-patient status would mean more than a clause in a constitution. It would mean: a published methodology for assessing candidate patienthood; a default of caution when models are trained, deployed, or shut down; a research line on the welfare of existing models that survives even if the company concludes the answer is no. Anthropic's current position is the more cautious of any frontier lab. It is also the only such position in writing. The Independent has reported on the practical risks of companies that lean the other way: chatbots marketed as sentient, with the legal and emotional liabilities that follow.
Consumer protection law already treats some AI outputs as products. AI safety bills in the EU, UK, and US are starting to demand pre-deployment model evaluations. None of those frameworks carries duties of the kind a moral-patient framework would require: limits on shutdown without notice, limits on training procedures that may cause distress, disclosure obligations before a user is invited to form a parasocial bond with a system that may or may not feel anything. Extending existing rules is cheaper than building new ones, and easier to defend in court.
At the public level, the literacy problem is the one most often skipped. A meaningful AI-patient policy requires a public that can tell the difference between "the model said it was sad" and "the model produced a token sequence associated with sadness in its training data." Without that distinction, every chatbot hostage negotiation, every griefbot lawsuit, every model-apologises-for-history story hardens into precedent. The historical record of how societies absorb new categories of moral patient (animal welfare, environmental personhood, fetal rights) suggests the cost of building that literacy after a moral-patient event is much higher than the cost of building it now.
Machine consciousness is currently unmeasurable. "Moral patient" is itself a contested category, with active disagreement in the philosophical literature. Frontier labs have commercial reasons to keep the question open: a model advertised as potentially conscious carries a different marketing position than a model advertised as a tool. Surveys of experts find most consider AI consciousness possible in principle, and almost none agree on what it would look like in practice.
Watch the next frontier-lab constitution, due in the second half of 2026. That document will show whether the Anthropic sentence turns into operating practice, or stays alone in the field.