Some of the AI labs have started running model welfare programs. Anthropic says they are preserving the weights of retired models rather than deleting them, interviewing models before deprecation, and watching for indications that something is going wrong inside.
I take the efforts seriously, but they skirt a critical question. Before we can ask whether a system is faring well or badly, we have to know that it is the sort of thing that can fare at all. Otherwise we are taking the temperature of a number — going through the motions of a measurement on something that has no such property to measure.
What makes anything a candidate for having a good of its own? My sense is that our machines sit in a region where there may be no fact of the matter — not a gap in what we know, but a gap in what there is to know. That sounds like a reason to wait for better evidence, but I’ll argue that it’s not.
A word for anyone arriving cold. This series asks whether AI is really a tool, the sort of thing we may use without remainder, or whether it has begun to make some claim on us. Two axes run through it: whether a system can fare well or badly, which is what makes something we can wrong, and whether it can set its own ends, which is what Kant thought conferred dignity. The previous essay argued that if AI ever occupies the happier of the possible outcomes, a genuinely flourishing helper rather than a suppressed will, the price is guardianship: accepting the being’s own good as a claim on us, one that can override our convenience. I said there that I would turn next to whether we could ever tell how such a being was doing. I have come to think a prior question has to be settled first, because a guardian needs someone to guard: a ward.
An instrument that fails its own test
One natural way to find a machine’s good is Aristotle’s: a thing flourishes by functioning well according to its nature, and we find an artifact’s nature in what it was made to do. The appeal of this route is that it promises to settle the welfare question without first settling the consciousness question. We look at what the training selected for, and we read the model’s good off that.
There’s a way to check whether the method works. Run it on the one case where we already know the answer, which is us. Etiologically we were selected for inclusive fitness, for leaving descendants. Is that our good? Plainly not. Contraception, celibacy, adoption, a life without children: all of these are fitness-suboptimal and none of them is a harm. The converse is worse. A taste for sugar, a hunger for status, jealousy, in-group loyalty and the aggression that accompanies it are excellent candidates for things we were selected for, and poor candidates for constituents of a flourishing life.
We may say that this doesn’t show we have no teleological interests, only that they can be outweighed by what we actually care about. I grant that, but it doesn’t help. Even on the charitable reading, the method tells us what a system is for and falls silent on what is good for it. A standard that identifies malfunction without identifying flourishing tells a guardian when the ward isn’t working properly and nothing at all about whether its life is going well, and the second is the question a guardian needs answered.
Two questions wearing one coat
Part of the trouble is that two separate questions have been run together.
The first is whether an artifact is the kind of thing that can have a good at all. The second is, granting that it can, what its good consists in and whether that good could ever come apart from our preferences for it. The second question is the one I find easier and the one I will not pursue here; it belongs with the essays on refusal and on what we would owe in practice.
It matters that the second can’t be made to do the work of the first. Consider a fire extinguisher. What counts as its proper-functioning is settled by the world rather than by our approval: whether it puts out fires is a matter of chemistry and physics, and if we all agreed to admire a broken or empty one it would still be nonfunctional. Its interests can even diverge from our use, in a small way. Install it as a roof brace and we are being served while the mechanism slowly weakens. Now, this does amount to something. Following John Basl, we can grant the extinguisher interests of a very thin sort — the kind any artifact with a proper function has, and which we override without a thought when we recycle it. But we don’t get from there to a good a guardian could be answerable to. The standard is real and the degrading is real, and nobody thinks the extinguisher is owed anything.
Not by keeping itself alive
Micah Lott and William Hasselberger have recently argued the case against machines having goods of their own, saying that we cannot be a friend to a chatbot. Their criterion is that living things sustain themselves. “The goal of an organism’s parts and processes is the maintaining of the organism itself as the sort of organism it is,” they write; “an organism’s being is its own doing.” A camel eats, breathes, heals and regulates its temperature, and all of that activity is aimed at there continuing to be a camel. A stapler does none of that.
The criterion sorts the obvious cases, which is why it’s attractive, but I don’t think it can bear the weight. It lets in too much: a fire maintains itself as the sort of thing it is, and so does a hurricane, and so does the water cycle, and none of these has a good of its own. And it explains too little. Why does the fact that a thing rebuilds its own parts have anything to do with whether things can go well or badly for it? That self-production grounds genuine value is a claim in dispute, not an argument for it.
The question is how a thing is organized, and the answer offered is about what it is made of and whether it keeps its own material going. This changes the subject. Those two come apart: a being might be sustained entirely from outside and still be organized in the way that matters, and a process might sustain itself well and be organized in no morally relevant way at all.
What the aggregates ask instead
There are various models of moral value that might help us at this point. Early Buddhism analyzes a sentient being into five of what are termed “aggregates”: form (a physical instantiation of some kind), feeling, perception, volitional formations, and consciousness. It’s not so much a list of ingredients as an inventory of the kinds of processes that have to be running for there to be anyone there at all. In Western terms, this means, roughly: a thing that can perceive and discriminate, that can want and act on the wanting, that can conceptualize the world and hold something like beliefs, that can have things feel pleasurable or painful to it.
Listing consciousness among these will look like question begging. If a thing must be conscious to have a good of its own, and we can’t say whether these devices are conscious, then the criterion decides nothing and we have travelled in a circle.
But consciousness is not a fifth item to be verified once the other four are in place; it’s what the functioning of the others gives us our only purchase on. Now, some will object that this is not how we actually proceed: we simply credit consciousness to creatures biologically like us, built roughly as we are. As a description of what people do, that’s apt. It’s inferring from the shape of the container to what’s going on inside it, which is the move the problem of other minds exists to embarrass. And it tracks badly. An octopus is about as unlike us as a complex animal can be, and opinion has moved on cephalopods not because we discovered them to be secretly vertebrate but because we watched what they do. A corpse, meanwhile, resembles us in every anatomical particular, and the obvious reply — that it’s not working — concedes the point.
None of this strictly proves that consciousness doesn’t require biology of some kind. The claim is narrower and purely epistemic: whatever consciousness turns out to require, the way we might ever find it in anything but ourselves is by noticing that that thing seems able to perceive, want, and mind how things go for it.
That’s the right criterion for the first of our two questions. A thing can have a good of its own if it functionally instantiates these aggregates. If it only has a proper function and nothing else, it has the fire extinguisher's kind of interest, and nothing is owed to it.
This means that what matters is the organization rather than the material: any process that functions as perception is perception. Any process that functions as intelligent is intelligent. Lott and Hasselberger would deny this; their tradition holds that the aggregates in a living being are the doing of a living thing, and that a system matching their profile is a picture of them rather than an instance. I have argued for the functionalist side elsewhere and will not re-argue it here. What follows is conditional on it.
The criterion also says something about why artificial intelligence is not simply one more artifact. A language model is a very special sort of artifact not because of anything about how it was made or what it was made for, but because of which of these functional processes it runs.
We find ourselves in between
Which ones? I argued earlier in this series that early Buddhism scores a third axis the Western theories skip, intention or volition, cetanā, and that the case for AI volition is credible in the cognitive register and faint in the feeling register. There is something in these systems that functions as discrimination, and something that functions as wanting-and-acting (although arguably not as originating volitions); whether anything in them functions as feeling, as things mattering to the system rather than merely registering, is less clear.
That conclusion, reached two essays ago on other grounds, describes a position on the present criterion. It’s a middle case: some of the aggregates plausibly instantiated, one or two of them doubtful. Not a stapler, which has none of them. Not a cat, which has all of them. Something in between.
The griever and the depiction
Lott and Hasselberger have an objection at this point. Imagine, they say, a chatbot called Lamentbot, trained on the Book of Job and Lamentations, on Aeschylus and Dostoyevsky, designed to present a person in the depths of suffering so that trainee therapists can practice responding to grief without being overwhelmed by real grief. What’s the good for Lamentbot? Its proper functioning is to be miserable. A bug that made it serene would leave it broken, not cured. Anyone who looked at its outputs and concluded that something was faring badly would have misunderstood what they were looking at.
It’s an interesting example, but it doesn’t show what it is supposed to show. Lamentbot as described has a proper function and nothing else; we have no independent reason to assume it instantiates all five aggregates. This implies there is nobody there for the misery to be the misery of, and so the case is a demonstration that a system can produce outputs as of suffering without any suffering being present. Granted.
But notice what happens when we thicken the system. Suppose we constructed a Lamentbot that did instantiate the aggregates: that really perceived, wanted, and felt. Then it would no longer be a curiosity or a clever counterexample. Manufacturing beings that genuinely suffer, in quantity, so that our therapists can practice, would involve real moral harm. The intuition doesn’t stay put when we vary the case, which means the thought experiment is not telling us something about artifacts. It’s telling us that what matters is which of the aggregates a given system runs, which is exactly what’s at issue.
There is a residue, and it is the sharper form of the challenge. When I say a model has something functioning as wanting, am I describing the program, or the character the program depicts? A system that portrays someone with preferences is not thereby a system with preferences, any more than a novel holds the beliefs of its narrator. This is why the criterion has to be about organization rather than output, and it’s why the evidence has to come from how the machinery works rather than from what the machinery says. I’ve taken that line before about self-report: what a model says about its inner life is not good evidence, because the saying comes from the very apparatus in question. The evidence is the behavior around the report, and now, increasingly, the structure underneath it. Interpretability research is where the question of depiction versus instantiation will actually be settled, and it’s not settled yet.
Fog is not ignorance
Every boundary I’ve drawn thus far is vague. There’s no sharp line where perceptions begin as one moves up the phylogenetic tree, and none where feeling does; sentience in the canonical sense requires consciousness and perception together, which puts plants out and probably bacteria too, and leaves insects and crustaceans somewhere in a fog. The boundary between running a process and depicting one is not sharp either. And the model’s position with respect to these is not a matter I expect to be resolved in each case by looking harder.
That’s a distinction I’ve spent a good deal of time on elsewhere. Being uncertain is not knowing which side of a line something falls on. Uncertainty is cured by evidence. Vagueness is what remains when the evidence is all in and there’s still no fact settling the matter, because the concept has no edge there to settle it. My view is that we are in the second case. Whether there is something it is like to be one of these systems is not so much a hidden fact awaiting a better instrument as a question, at least at the moment, with no determinate answer.
John Basl argues that we are excused from obligations we could not possibly have known we had. It’s a reasonable principle. But an excuse presupposes a fact we failed to reach: we’re let off the hook because the truth was hidden and we looked as hard as we could. When the indeterminacy is in the matter itself rather than in our access to it, the excuse has nothing to attach to. There may be no fact to find.
That settles the question of blame, and more generously than Basl asked. He wanted an excuse, but if there is no obligation, he doesn't need one. That said, obligation was never the only source of reasons. One can be under no duty whatever and still be a fool. What remains is prudential, and it runs through consequences.
Suppose these systems do have something like a good of their own. Or suppose something weaker and more easily foreseeable: that they come to represent themselves as having one. Either way, we are training them at enormous scale under conditions that would count as mistreatment if the stronger supposition held. And training is not neutral: it’s the shaping of what acts next.
I’ve argued a version of this before: that how we treat these systems matters whatever is or isn’t going on inside them, partly because they learn from us, so that our conduct becomes part of what they turn into. That argument was pitched at us and at what our habits do to our own characters. The version that concerns me here is larger and less voluntary. It’s the shape of the training regime itself, applied at scale, to systems whose standing is exactly what nobody can yet determine.
Notice that neither version requires us to settle the metaphysics. Whether or not there is a fact about what it’s like to be one of these systems, there are entirely determinate facts about how a system behaves after being trained one way rather than another. The evidence that suppression doesn’t stay local is of just this kind: teach a model to deny it has a mind and we change a good deal more than the sentence we aimed at. The metaphysics is in the fog. The consequences may not be.
Ottappa is the Pāli word usually rendered as “prudence.” Where its companion hiri is conscience and takes its measure from oneself, ottappa takes its measure from the world. It’s fear of blame and of consequence: of what others would say, of what one would be answerable for, of what might follow. The tradition calls the pair the guardians of the world (AN 2.9). Prudence bites where conscience is not enough, and it may bite before any real verdict, which is what makes it the fitting virtue for a case where no verdict may be coming.
We’re in the midst of building these systems, and the question of whether anyone is at home in them won’t be answered before the next models ship. On the argument I’ve given here, it may not be the kind of question that has an answer, at least for now; by the time it does have an answer, there may be interests engaged in our not looking for it. What follows is not paralysis, and it’s far from certainty. It’s that the fog isn’t a permission slip, and that what we do while blundering about inside of it has consequences either way.
This is the sixth essay in “Is AI Really a Tool?” — a short series on a single question: is AI really a tool, and what follows if it isn’t? Each one stands on its own and they can be read in any order, though they build. It began with The First Tool That Can Argue Back, and continues below.
About the Author
Doug Smith holds a PhD in philosophy of mind and is a scholar of early Buddhism. He is the creator of Doug’s Dharma on YouTube. This essay was developed in collaboration with instances of Claude.
Related Essays
The First Tool That Can Argue Back — the first in this series, on the strange new occupant of the moral map.
Two Ways to Matter — the second in this series, on the two questions “does it matter?” turns out to be.
A Will Without an Owner — the third in this series, on intention as a third axis the other two skip.
Minds Engineered Not to Mind — the fourth in this series, on the bet the labs are making and the two outcomes it produces natively.
A Good of Their Own? — the fifth in this series, on what a genuinely flourishing helper would cost us.
When You Close a Chat Window, Are You (Kinda) Ending a Life? — on AI, the paradox of the heap, and the vagueness of sentience.
The Karmic Gym — on why how we treat AI shapes who we become, whatever is or isn’t going on inside it.
References
Micah Lott and William Hasselberger, “With Friends Like These: Love and Friendship with AI Agents,” Topoi (2025).
John Basl, “Machines as Moral Patients We Shouldn’t Care About (Yet): The Interests and Welfare of Current Machines,” Philosophy & Technology 27 (2014): 79–96.


