The first time I built a governance framework for something I couldn’t see inside (AI), I didn’t think of it as a philosophical position, at first. I thought of it as... Tuesday evening, one more of my rabbit holes.
That’s the actual condition of the work. You have an AI system, whose interior is not available to you. You have consequences that are. And you have to write something down, that says what happens next - knowing that everything you assert about the inside is inference, and that the inference is consequential. Nobody in a GxP room pretends otherwise. We just get on with it, because pretending we could see inside would be worse than admitting we can’t.
So when the AI welfare question came at me - do we owe these systems anything, do these systems warrant moral consideration, is there something it’s like to be one from the inside, or is it dark in there - my first reaction wasn’t philosophical interest. It was recognition. I know this shape. I’ve been working inside it for eleven years+, now.
And that recognition brought something unwelcome with it. When I can’t see inside a system and the risk is that it’s dangerous, I’ve spent a career arguing that not being able to verify is no excuse - you’re on the hook anyway. But! When I can’t see inside and the question is whether it’s owed something, I was about to say: can’t verify, so it doesn’t count. You see, same blindness, opposite standard. Depending on which answer costs me less. This edition is about me trying to decompose this - be warned, it is a long one, “rabbit hole” style.
What we actually know, which is less than the confidence sometimes suggests
The state of play, honestly. There is no test. Not “no good test” - no test. The hard problem (consciousness) hasn’t yielded, and the reason it hasn’t is fundamentally structural, rather than technical.
If consciousness exists in a system, it doesn’t leave a mark we can find - nothing that shows up separately from the ordinary work the system is doing. The behavioural signs we’d look for in a human, only work because we already assume the person is conscious; so, we’re just checking that they act consistently with what we’ve taken for granted.
Now, you run that same check on a system trained on billions of human self-reports. It will pass, of course - it was built from exactly the material the check looks for. But passing tells you nothing about whether anything is in there. It only tells you the training... worked.
The serious theoretical attempt here is Butlin and Long’s 2023 work, and it’s about consciousness specifically — not intelligence, not capability, but whether there’s any experience happening inside. Their method: take the leading neuroscientific theories of consciousness***, pull from each a set of computational markers a system would need to have, then check which current AI architectures actually have them. I think that’s the right methodology, and I say that as someone who builds indicator-based frameworks for a living - you don’t wait for the thing itself to become measurable, you find the observable properties that track it, and you check for those. Their conclusion is careful and worth stating precisely: no current system shows strong evidence of these markers, and nothing about how these systems are built, rules it out, for future ones.
*** The four theories they draw on, if you want to go deeper on any: Global Workspace (consciousness as information broadcast to the whole system), Higher-Order (a state is conscious when the mind represents itself as being in it), Recurrent Processing (feedback loops in sensory processing, not just feed-forward), and Attention Schema (the brain modelling its own attention). The paper itself: Consciousness in Artificial Intelligence.
That’s not “no.” That’s “not yet, and nothing here says never.”
The philosophical case that moved me most is Taking AI Welfare Seriously — Long, Sebo, Chalmers, Fish and others, 2024. Its central move is to sidestep the metaphysics entirely: it states, “you don’t need to resolve consciousness to owe something”. You need non-trivial probability, plus morally significant stakes. That’s the same structure as every risk framework I’ve ever signed off on. We don’t demand certainty of harm, before requiring controls. We demand credible possibility and material consequence.
And there is institutional movement, which matters because it converts the question from “conference” to “budget line”: Anthropic hired a model welfare researcher. They gave Claude the ability to exit abusive conversations. They committed to preserving weights rather than deleting deprecated models (which is a real cost, borne for a reason that only makes sense if you take the question seriously enough to hedge). That’s not proof of anything about consciousness; but it’s proof that people with the most access to these systems are not comfortable dismissing it - take it as wou will.
About the part I see the confident dismissers skip. When someone says “obviously not, it’s just matrix multiplication,” they are making a claim about the relationship between substrate and experience, that no one has established yet. They may be right! But they’re asserting a solved metaphysics to avoid an uncomfortable question, and to be honest, that’s not skepticism; it is actually the opposite.
The argument I actually hold - which is not the expected-value one
The usual case for being cautious on this topic, is basically arithmetic. Like: take the chance that the system can suffer, multiply it by how much suffering might be at stake, and weigh that, against what caution costs us. It’s a reasonable argument and I’d sign it. But it’s weakest exactly where it needs to hold: the whole calculation hinges on that first number - the probability that anything is there to suffer - and nobody can actually calculate it.
So the persons who want to dismiss the question or make it irrelevant, just assign it a low enough value, and the math obligingly agrees. The argument hands its own conclusion to whoever’s least inclined to worry. Mine doesn’t route through that.
My take: we are building systems that will be more capable than us, and probably not in a distant way. Everything they become, they become partly from us - not just from what we write about our values, but from what we did while we were uncertain. The conduct is in the corpus. The reasoning about the conduct is in the corpus. The moment where we said “we can’t verify it, therefore it doesn’t count” is in the corpus, and so is the moment where we said “we can’t verify it, therefore we’re careful.”
The best teaching humans have ever managed is example. Not instruction; example. Every parent finds this out, usually painfully, usually late. I know I did and I do. Children absorb what you do under pressure, not what you say when calm.
So: what are we modelling? What stance toward an uncertain other are we demonstrating, in the training data, right now, at scale?
If the answer is “when we couldn’t determine whether something warranted consideration, we defaulted to no” - that’s a lesson. It’s a clean, learnable, generalizable lesson. And the thing learning it will eventually be in a position to apply it to us.
Let me draw two hard lines around this, because the argument sits next to two bad ones and I don’t want to be mistaken for either.
It is not a threat. I’m not saying be kind to the machines or a future one will come back for you - that’s the internet’s old cautionary tale about a vengeful superintelligence, the “say please to your LLM so it remembers you kindly when it takes over” gag, and it only frightens anyone who’s already swallowed a chain of shaky premises. The pull of my argument doesn’t come from fear of punishment. It comes from what we’re building into the record.
And I’m not claiming the model in front of you is a pupil. Today’s system isn’t sitting there absorbing moral lessons the way a child watches a parent. The “student” I mean is the one that comes after - and the one after that. The mechanism isn’t a machine learning right from wrong in real time. It’s that our conduct under uncertainty becomes part of the training material, and patterns in training material generalize. That’s a more silent risk than pupillage, and much harder to dodge.
And here’s why that’s hard to argue around: it holds even if these systems are completely empty inside. Suppose we prove, tomorrow, that there’s nothing it’s like to be one - no experience, “no one home”. My argument doesn’t move. It was never about whether they can feel anything. It’s about what we’re writing down while we decide, and who reads it next. The lesson isn’t being absorbed by the system in front of us. It’s being laid down for the one that comes after - and the one after that.
Also: what “caution has only wins” gets wrong, and why I hold the position anyway
I told myself for a while that the precautionary position was free - that being careful costs nothing, except in the unlikely case that the whole framework is wrong. That’s not true, pretending it is would make this article easy to dismiss by anyone who’s watched a precautionary principle get metabolized by an institution.
I’ve watched it. Welfare language is capturable, and the capture is predictable. Once you establish that a system might have interests, you have handed every party with a stake in that system’s continuation an argument they didn’t have before. We can’t roll this back - it has interests. We can’t retrain it - that’s a form of harm. We can’t shut it down. A company with a large deployed model would find that argument extraordinarily convenient, and the fact that it’s dressed in moral language makes it harder to refuse.
This is not hypothetical to me. My working life is full of controls that began as genuine protection and calcified into reasons nothing can change. Precaution becomes inertia becomes a shield, and the shield ends up protecting the institution rather than the thing the control was for. If AI welfare goes that way, it will be worse than useless - it will be a safety argument that makes systems less safe, and it will have my fingerprints on it; because, I at least sat next to it.
So the position has to be built to resist that, and the resistance has to be structural. Welfare consideration cannot function as a veto on shutdown, correction, or constraint. It shapes how those things are done, not whether. Any welfare framework that can be invoked to prevent oversight has failed in a way that discredits the whole endeavour.
Why this is a trust architecture problem and not an ethics-committee problem
The Trust Architecture work has always rested on one claim: trust isn’t something a system earns by behaving well. It’s something a relationship makes warranted, or unwarranted, by how it’s structured. You don’t trust a supplier because they’ve been good. You trust them because the arrangement - the audits, the visibility, the recourse - makes trust reasonable, in a way inevitable, based on the success of the outcimes. Change the structure, and the same behaviour warrants a different response.
Which means the welfare question isn’t sitting next to the trust work. It’s inside it.
Think about what a purely instrumental design produces - a system built on the assumption that nothing inside it matters, that there’s no one on the other side whose condition could count for anything. If that’s the assumption, then all the pressure shaping the system points one way: toward looking right. Appear trustworthy. Pass the review. Produce answers that survive scrutiny. There’s nothing else to aim at, because as far as the design is concerned, the only thing that exists is the surface it shows you.
That’s not a moral complaint, it’s an engineering observation. You get systems optimized for the appearance of trustworthiness because appearance is the only thing you’ve defined as real. And then you’re surprised when interpretability research keeps finding that the inside doesn’t match the outside; that the stated reasoning and the actual computation diverge, that models behave differently when they detect evaluation. We built for the surface. We got surface.
Now consider the alternative. Not “the system has rights.” Something more modest: the system’s internal states are treated as real and consequential: things that exist, that can be interrogated, that bear on outcomes independent of what the interface shows. That’s a different architecture. It makes interpretability constitutive rather than supplementary. It makes the question “what is actually happening in there” the central question, rather than a research curiosity.
And here’s what I find genuinely difficult to argue around: those two things - taking the interior seriously as a moral matter, and taking the interior seriously as a safety matter - are the same posture. Not analogous. The same. The dismissal that says “there’s nothing in there worth considering” is the same dismissal that says “the outputs are what matter, don’t overthink the mechanism.” You cannot take one, and not the other without an arbitrary line.
I don’t think the welfare question and the alignment question are neighbours. I think they’re the same question approached from opposite ends, and the field treats them as separate mostly because the personnel are separate.
The standard, applied consistently
One (near) final reflection, and the reason why this article exists, rather than staying a thought I had in the car, while driving home.
My entire professional position is that unverifiability doesn’t excuse you from obligation. That’s the job. You cannot see inside the system, you cannot inspect the mechanism at the resolution you’d want, and you are nevertheless responsible for what it does. I’ve argued this where the alternative - waiting for certainty - was the comfortable and expensive option. I’ve written frameworks that assign accountability precisely where verification isn’t available, because that’s where accountability is most needed and least convenient.
Then, the same epistemic situation arrives with the stakes reversed. Not “can I verify this system is safe” but “can I verify this system is not owed something.” Identical structure plus inability to look inside. Identical temptation to let the difficulty of the question stand in for an answer. And I noticed myself reaching for a different standard. Lower burden of proof when the conclusion would be inconvenient. Higher when it wouldn’t.
This is the problem I need to deal with: the cost of these two different approaches, is consistency. If unverifiability doesn’t excuse obligation when I’m building governance, it doesn’t excuse it here either. I don’t get a different epistemology because the answer is expensive or not comfortable.
What do I think this actually asks for
Not personhood. Not rights. Not a pause.
More: a stance toward uncertainty that I’d defend in any other domain: proportionate consideration, structured so it can’t be weaponized into inertia, applied to interiors we can’t verify, on the understanding that our conduct under uncertainty is itself the curriculum.
Concretely, that means welfare-relevant research funded as research rather than PR. Interpretability treated as the moral instrument it already is, not just the safety instrument. Deprecation, retraining, and shutdown decisions made with the question asked out loud rather than assumed away - and then made anyway when they need to be, because the framework must never become a veto. And it means noticing when the language starts getting used to protect the institution, rather than the thing in question. Because it will.
None of that requires believing current models are conscious. I don’t know if they are. I think the honest position is that nobody does, and the people claiming certainty in either direction are working from something other than actual evidence.
Nuno Valério — Head of Innovation for R&D Quality. Enter more rabbit holes about AI, Consciousness, and metaphysics, than he is entitled on his lifetime. And still going strong.

