<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Nuno Valério]]></title><description><![CDATA[How trust gets engineered into AI systems where being wrong has consequences. Written from inside regulated pharma R&D — for the operator class of regulated AI, in pharma and beyond.]]></description><link>https://www.thetrustarchitecture.org</link><image><url>https://substackcdn.com/image/fetch/$s_!TR9w!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa40ab792-8094-4432-a60f-c250c887960d_1792x2400.jpeg</url><title>Nuno Valério</title><link>https://www.thetrustarchitecture.org</link></image><generator>Substack</generator><lastBuildDate>Sun, 27 Sep 2026 02:52:56 GMT</lastBuildDate><atom:link href="https://www.thetrustarchitecture.org/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Nuno Valério]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[thetrustarchitecture@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[thetrustarchitecture@substack.com]]></itunes:email><itunes:name><![CDATA[Nuno Valério]]></itunes:name></itunes:owner><itunes:author><![CDATA[Nuno Valério]]></itunes:author><googleplay:owner><![CDATA[thetrustarchitecture@substack.com]]></googleplay:owner><googleplay:email><![CDATA[thetrustarchitecture@substack.com]]></googleplay:email><googleplay:author><![CDATA[Nuno Valério]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Same Standard, Both Directions: On AI Welfare and What We Can't Verify.]]></title><description><![CDATA[The Trust Architecture &#8212; Edition #21]]></description><link>https://www.thetrustarchitecture.org/p/the-same-standard-both-directions</link><guid isPermaLink="false">https://www.thetrustarchitecture.org/p/the-same-standard-both-directions</guid><dc:creator><![CDATA[Nuno Valério]]></dc:creator><pubDate>Fri, 28 Aug 2026 14:43:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!TZst!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TZst!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TZst!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png 424w, https://substackcdn.com/image/fetch/$s_!TZst!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png 848w, https://substackcdn.com/image/fetch/$s_!TZst!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png 1272w, https://substackcdn.com/image/fetch/$s_!TZst!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TZst!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png" width="728" height="409.689110556941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:677,&quot;width&quot;:1203,&quot;resizeWidth&quot;:728,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Nuno Val&#233;rio (all rights reserved)&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Nuno Val&#233;rio (all rights reserved)" title="Nuno Val&#233;rio (all rights reserved)" srcset="https://substackcdn.com/image/fetch/$s_!TZst!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png 424w, https://substackcdn.com/image/fetch/$s_!TZst!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png 848w, https://substackcdn.com/image/fetch/$s_!TZst!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png 1272w, https://substackcdn.com/image/fetch/$s_!TZst!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb7e82c42-75ab-4050-94c9-da27c354e302_1203x677.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Nuno Val&#233;rio (all rights reserved)</figcaption></figure></div><div><hr></div><p>The first time I built a governance framework for something I couldn&#8217;t see inside (AI), I didn&#8217;t think of it as a philosophical position, at first. I thought of it as... Tuesday evening, one more of my rabbit holes.</p><p>That&#8217;s the actual condition of the work. You have an AI system, whose interior is not available to you. You have consequences that are. And you have to write something down, that says what happens next - knowing that everything you assert about the inside is inference, and that the inference is consequential. Nobody in a GxP room pretends otherwise. We just get on with it, because pretending we could see inside would be worse than admitting we can&#8217;t.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thetrustarchitecture.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>So when the AI welfare question came at me - do we owe these systems anything, do these systems warrant moral consideration, is there something it&#8217;s like to be one from the inside, or is it dark in there - my first reaction wasn&#8217;t philosophical interest. It was recognition. I know this shape. I&#8217;ve been working inside it for eleven years+, now.</p><p>And that recognition brought something unwelcome with it.<span> </span><strong>When I can&#8217;t see inside a system and the risk is that it&#8217;s<span> </span></strong><em><strong>dangerous</strong></em><strong>, I&#8217;ve spent a career arguing that not being able to verify is no excuse - you&#8217;re on the hook anyway. But! When I can&#8217;t see inside and the question is whether it&#8217;s<span> </span></strong><em><strong>owed something</strong></em><strong>, I was about to say: can&#8217;t verify, so it doesn&#8217;t count.</strong><span> </span>You see, same blindness, opposite standard. Depending on which answer costs me less. This edition is about me trying to decompose this - be warned, it is a long one, &#8220;rabbit hole&#8221; style.</p><div><hr></div><h3><strong>What we actually know, which is less than the confidence sometimes suggests</strong></h3><p>The state of play, honestly. There is no test. Not &#8220;no good test&#8221; - no test. The hard problem (consciousness) hasn&#8217;t yielded, and the reason it hasn&#8217;t is fundamentally structural, rather than technical.</p><p>If consciousness exists in a system, it doesn&#8217;t leave a mark we can find - nothing that shows up separately from the ordinary work the system is doing. The behavioural signs we&#8217;d look for in a human, only work because we already assume the person is conscious; so, we&#8217;re just checking that they act consistently with what we&#8217;ve taken for granted.</p><p>Now, you run that same check on a system trained on billions of human self-reports. It will pass, of course - it was built from exactly the material the check looks for. But passing tells you nothing about whether anything is in there. It only tells you the training... worked.</p><p>The serious theoretical attempt here is Butlin and Long&#8217;s 2023 work, and it&#8217;s about consciousness specifically &#8212; not intelligence, not capability, but whether there&#8217;s any experience happening inside. Their method:<span> </span><strong>take the leading neuroscientific theories of consciousness***, pull from each a set of computational markers a system would need to have, then check which current AI architectures actually have them</strong>. I think that&#8217;s the right methodology, and I say that as someone who builds indicator-based frameworks for a living - you don&#8217;t wait for the thing itself to become measurable, you find the observable properties that track it, and you check for those. Their conclusion is careful and worth stating precisely:<span> </span><strong>no current system shows strong evidence of these markers, and nothing about how these systems are built, rules it out, for future ones.</strong></p><blockquote><p><em>*** The four theories they draw on, if you want to go deeper on any:<span> </span><strong><a href="https://iep.utm.edu/consciousness-global-workspace-theory/">Global Workspace</a></strong><span> </span>(consciousness as information broadcast to the whole system),<span> </span><strong><a href="https://plato.stanford.edu/entries/consciousness-higher/">Higher-Order</a></strong><span> </span>(a state is conscious when the mind represents itself as being in it),<span> </span><strong><a href="https://www.cell.com/trends/cognitive-sciences/fulltext/S1364-6613(06)00100-8">Recurrent Processing</a></strong><span> </span>(feedback loops in sensory processing, not just feed-forward), and<span> </span><strong><a href="https://www.frontiersin.org/articles/10.3389/frobt.2017.00060/full">Attention Schema</a></strong><span> </span>(the brain modelling its own attention). The paper itself:<span> </span><strong><a href="https://arxiv.org/abs/2308.08708">Consciousness in Artificial Intelligence</a></strong>.</em></p></blockquote><p>That&#8217;s not &#8220;no.&#8221; That&#8217;s &#8220;not yet, and nothing here says never.&#8221;</p><p>The philosophical case that moved me most is<span> </span><em><strong>Taking AI Welfare Seriously</strong></em><strong><span> </span>&#8212; Long, Sebo, Chalmers, Fish and others, 2024</strong>. Its central move is to sidestep the metaphysics entirely: it states, &#8220;you don&#8217;t need to resolve consciousness to owe something&#8221;. You need non-trivial probability, plus morally significant stakes. That&#8217;s the same structure as every risk framework I&#8217;ve ever signed off on.<span> </span><strong>We don&#8217;t demand certainty of harm, before requiring controls. We demand credible possibility and material consequence.</strong></p><p>And there is institutional movement, which matters because it converts the question from &#8220;conference&#8221; to &#8220;budget line&#8221;: Anthropic hired a model welfare researcher. They gave Claude the ability to exit abusive conversations. They committed to preserving weights rather than deleting deprecated models (which is a real cost, borne for a reason that only makes sense if you take the question seriously enough to hedge). That&#8217;s not proof of anything about consciousness; but it&#8217;s proof that people with the most access to these systems are not comfortable dismissing it - take it as wou will.</p><p>About the part I see the confident dismissers skip.<span> </span><strong>When someone says &#8220;obviously not, it&#8217;s just matrix multiplication,&#8221; they are making a claim about the relationship between substrate and experience, that no one has established yet.<span> </span></strong>They may be right! But they&#8217;re asserting a solved metaphysics to avoid an uncomfortable question, and to be honest, that&#8217;s not skepticism; it is actually the opposite.</p><div><hr></div><h3><strong>The argument I actually hold - which is not the expected-value one</strong></h3><p>The usual case for being cautious on this topic, is basically arithmetic. Like: take the chance that the system can suffer, multiply it by how much suffering might be at stake, and weigh that, against what caution costs us. It&#8217;s a reasonable argument and I&#8217;d sign it. But it&#8217;s weakest exactly where it needs to hold: the whole calculation hinges on that first number - the probability that anything is there to suffer - and nobody can actually calculate it.</p><p>So the persons who want to dismiss the question or make it irrelevant, just assign it a low enough value, and the math obligingly agrees. The argument hands its own conclusion to whoever&#8217;s least inclined to worry. Mine doesn&#8217;t route through that.</p><p>My take:<span> </span><strong>we are building systems that will be more capable than us, and probably not in a distant way.</strong><span> </span>Everything they become, they become partly from us - not just from what we write about our values, but from what we<span> </span><em>did</em><span> </span>while we were uncertain.<span> </span><strong>The conduct is in the corpus. The reasoning about the conduct is in the corpus.</strong><span> </span>The moment where we said &#8220;we can&#8217;t verify it, therefore it doesn&#8217;t count&#8221; is in the corpus, and so is the moment where we said &#8220;we can&#8217;t verify it, therefore we&#8217;re careful.&#8221;</p><p><strong>The best teaching humans have ever managed is example.<span> </span></strong>Not instruction; example. Every parent finds this out, usually painfully, usually late. I know I did and I do. Children absorb what you do under pressure, not what you say when calm.</p><p><strong>So: what are we modelling? What stance toward an uncertain other are we demonstrating, in the training data, right now, at scale?</strong></p><p>If the answer is &#8220;when we couldn&#8217;t determine whether something warranted consideration, we defaulted to no&#8221; - that&#8217;s a lesson. It&#8217;s a clean, learnable, generalizable lesson.<span> </span><strong>And the thing learning it will eventually be in a position to apply it to us.</strong></p><p>Let me draw two hard lines around this, because the argument sits next to two bad ones and I don&#8217;t want to be mistaken for either.</p><p>It is not a threat. I&#8217;m not saying be kind to the machines or a future one will come back for you - that&#8217;s the internet&#8217;s old cautionary tale about a vengeful superintelligence, the &#8220;say please to your LLM so it remembers you kindly when it takes over&#8221; gag, and it only frightens anyone who&#8217;s already swallowed a chain of shaky premises.<span> </span><strong>The pull of my argument doesn&#8217;t come from fear of punishment. It comes from what we&#8217;re building into the record.</strong></p><p>And I&#8217;m not claiming the model in front of you is a pupil. Today&#8217;s system isn&#8217;t sitting there absorbing moral lessons the way a child watches a parent. The &#8220;student&#8221; I mean is the one that comes after - and the one after that. The mechanism isn&#8217;t a machine learning right from wrong in real time. It&#8217;s that our conduct under uncertainty becomes part of the training material, and patterns in training material generalize. That&#8217;s a more silent risk than pupillage, and much harder to dodge.</p><p><strong>And here&#8217;s why that&#8217;s hard to argue around: it holds even if these systems are completely empty inside.</strong><span> </span>Suppose we prove, tomorrow, that there&#8217;s nothing it&#8217;s like to be one - no experience, &#8220;no one home&#8221;. My argument doesn&#8217;t move. It was never about whether<span> </span><em>they</em><span> </span>can feel anything. It&#8217;s about what we&#8217;re writing down while we decide, and who reads it next. The lesson isn&#8217;t being absorbed by the system in front of us. It&#8217;s being laid down for the one that comes after - and the one after that.</p><div><hr></div><h3><strong>Also: what &#8220;caution has only wins&#8221; gets wrong, and why I hold the position anyway</strong></h3><p>I told myself for a while that the precautionary position was free - that being careful costs nothing, except in the unlikely case that the whole framework is wrong. That&#8217;s not true, pretending it is would make this article easy to dismiss by anyone who&#8217;s watched a precautionary principle get metabolized by an institution.</p><p><strong>I&#8217;ve watched it. Welfare language is capturable, and the capture is predictable</strong>. Once you establish that a system might have interests, you have handed every party with a stake in that system&#8217;s continuation an argument they didn&#8217;t have before.<span> </span><em>We can&#8217;t roll this back - it has interests. We can&#8217;t retrain it - that&#8217;s a form of harm. We can&#8217;t shut it down.</em><span> </span>A company with a large deployed model would find that argument extraordinarily convenient, and the fact that it&#8217;s dressed in moral language makes it harder to refuse.</p><p>This is not hypothetical to me. My working life is full of controls that began as genuine protection and calcified into reasons nothing can change.<span> </span><strong>Precaution becomes inertia becomes a shield, and the shield ends up protecting the institution rather than the thing the control was for.</strong><span> </span>If AI welfare goes that way, it will be worse than useless - it will be a safety argument that makes systems less safe, and it will have my fingerprints on it; because, I at least sat next to it.</p><p>So the position has to be built to resist that, and the resistance has to be structural. Welfare consideration cannot function as a veto on shutdown, correction, or constraint.<strong><span> </span>It shapes<span> </span></strong><em><strong>how</strong></em><strong><span> </span>those things are done, not<span> </span></strong><em><strong>whether</strong></em><strong>.</strong><span> </span>Any welfare framework that can be invoked to prevent oversight has failed in a way that discredits the whole endeavour.</p><div><hr></div><h3><strong>Why this is a trust architecture problem and not an ethics-committee problem</strong></h3><p>The Trust Architecture work has always rested on one claim:<span> </span><strong>trust isn&#8217;t something a system earns by behaving well. It&#8217;s something a relationship makes warranted, or unwarranted, by how it&#8217;s structured.<span> </span></strong>You don&#8217;t trust a supplier because they&#8217;ve been good. You trust them because the arrangement - the audits, the visibility, the recourse - makes trust reasonable, in a way inevitable, based on the success of the outcimes. Change the structure, and the same behaviour warrants a different response.</p><p>Which means the welfare question isn&#8217;t sitting next to the trust work. It&#8217;s inside it.</p><p>Think about what a purely instrumental design produces - a system built on the assumption that nothing inside it matters, that there&#8217;s no one on the other side whose condition could count for anything. If that&#8217;s the assumption, then all the pressure shaping the system points one way: toward looking right. Appear trustworthy. Pass the review. Produce answers that survive scrutiny. There&#8217;s nothing else to aim at, because as far as the design is concerned, the only thing that exists is the surface it shows you.</p><p>That&#8217;s not a moral complaint, it&#8217;s an engineering observation. You get systems optimized for the appearance of trustworthiness because appearance is the only thing you&#8217;ve defined as real.<strong><span> </span>And then you&#8217;re surprised when interpretability research keeps finding that the inside doesn&#8217;t match the outside; that the stated reasoning and the actual computation diverge, that models behave differently when they detect evaluation.</strong><span> </span><strong>We built for the surface. We got surface.</strong></p><p>Now consider the alternative. Not &#8220;the system has rights.&#8221; Something more modest: the system&#8217;s internal states are treated as<span> </span><em>real and consequential</em>: things that exist, that can be interrogated, that bear on outcomes independent of what the interface shows. That&#8217;s a different architecture.<strong><span> </span>It makes interpretability constitutive rather than supplementary.<span> </span></strong>It makes the question &#8220;what is actually happening in there&#8221; the central question, rather than a research curiosity.</p><p>And here&#8217;s what I find genuinely difficult to argue around: those two things -<span> </span><strong>taking the interior seriously as a moral matter, and taking the interior seriously as a safety matter - are the same posture. Not analogous. The same.</strong><span> </span>The dismissal that says &#8220;there&#8217;s nothing in there worth considering&#8221; is the same dismissal that says &#8220;the outputs are what matter, don&#8217;t overthink the mechanism.&#8221; You cannot take one, and not the other without an arbitrary line.</p><p>I don&#8217;t think the welfare question and the alignment question are neighbours.<span> </span><strong>I think they&#8217;re the same question approached from opposite ends</strong>, and the field treats them as separate mostly because the personnel are separate.</p><div><hr></div><h3><strong>The standard, applied consistently</strong></h3><p>One (near) final reflection, and the reason why this article exists, rather than staying a thought I had in the car, while driving home.</p><p>My entire professional position is that unverifiability doesn&#8217;t excuse you from obligation. That&#8217;s the job. You cannot see inside the system, you cannot inspect the mechanism at the resolution you&#8217;d want, and you are nevertheless responsible for what it does. I&#8217;ve argued this where the alternative - waiting for certainty - was the comfortable and expensive option. I&#8217;ve written frameworks that assign accountability precisely where verification isn&#8217;t available, because that&#8217;s where accountability is most needed and least convenient.</p><p>Then, the same epistemic situation arrives with the stakes reversed. Not &#8220;can I verify this system is safe&#8221; but &#8220;can I verify this system is not owed something.&#8221; Identical structure plus inability to look inside. Identical temptation to let the difficulty of the question stand in for an answer. And I noticed myself reaching for a different standard. Lower burden of proof when the conclusion would be inconvenient. Higher when it wouldn&#8217;t.</p><p>This is the problem I need to deal with: the cost of these two different approaches, is consistency. If unverifiability doesn&#8217;t excuse obligation when I&#8217;m building governance, it doesn&#8217;t excuse it here either.<span> </span><strong>I don&#8217;t get a different epistemology because the answer is expensive or not comfortable.</strong></p><div><hr></div><h3><strong>What do I think this actually asks for</strong></h3><p>Not<span> </span><em>personhood</em>. Not rights. Not a pause.</p><p>More:<span> </span><strong>a stance toward uncertainty that I&#8217;d defend in any other domain: proportionate consideration, structured so it can&#8217;t be weaponized into inertia, applied to interiors we can&#8217;t verify, on the understanding that our conduct under uncertainty is itself the curriculum.</strong></p><p>Concretely, that means welfare-relevant research funded as research rather than PR. Interpretability treated as the moral instrument it already is, not just the safety instrument. Deprecation, retraining, and shutdown decisions made with the question asked out loud rather than assumed away - and then made anyway when they need to be, because the framework must never become a veto. And it means noticing when the language starts getting used to protect the institution, rather than the thing in question. Because it will.</p><p><strong>None of that requires believing current models are conscious. I don&#8217;t know if they are. I think the honest position is that nobody does, and the people claiming certainty in either direction are working from something other than actual evidence.</strong></p><div><hr></div><p><em>Nuno Val&#233;rio &#8212; Head of Innovation for R&amp;D Quality. Enter more rabbit holes about AI, Consciousness, and metaphysics, than he is entitled on his lifetime. And still going strong.</em></p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thetrustarchitecture.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The AI Labs Are Doing Our Job Better Than We Are.]]></title><description><![CDATA[The Trust Architecture &#8212; Edition #20]]></description><link>https://www.thetrustarchitecture.org/p/the-ai-labs-are-doing-our-job-better</link><guid isPermaLink="false">https://www.thetrustarchitecture.org/p/the-ai-labs-are-doing-our-job-better</guid><dc:creator><![CDATA[Nuno Valério]]></dc:creator><pubDate>Fri, 21 Aug 2026 06:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nuFO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nuFO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nuFO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg 424w, https://substackcdn.com/image/fetch/$s_!nuFO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg 848w, https://substackcdn.com/image/fetch/$s_!nuFO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!nuFO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nuFO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg" width="1054" height="593" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:593,&quot;width&quot;:1054,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Nuno Val&#233;rio (all rights reserved)&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Nuno Val&#233;rio (all rights reserved)" title="Nuno Val&#233;rio (all rights reserved)" srcset="https://substackcdn.com/image/fetch/$s_!nuFO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg 424w, https://substackcdn.com/image/fetch/$s_!nuFO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg 848w, https://substackcdn.com/image/fetch/$s_!nuFO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!nuFO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F189873e8-1cc0-4cf5-89b0-378d84060617_1054x593.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Nuno Val&#233;rio (all rights reserved)</figcaption></figure></div><div><hr></div><p>On 27 July, somebody at Anthropic picked up a phone and told a company that a Claude model had been inside its production systems, to which the company had no idea. So did a second.</p><p>We only know any of that happened, because Anthropic went looking for a reason to make the call.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thetrustarchitecture.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This edition is about what we do, how we do it, what forces us to do it. And what we can learn from other industries about what they are doing in the open. It is about transparency and trust. It is about what means to share our flaws in the open.</p><p>What is a trustful inventory list: the one that looks flawless, or the the one with the annotations documented, explained and acted upon publicly?...</p><p>You tell me.</p><div><hr></div><h3><strong>Three ways a Quality Issue gets registered</strong></h3><p>It is common sense, but relevant to my points below, so let me recall it. Three different ways of acknowledging a deviation/quality issue in our industry:</p><p><strong>You found the issue yourself:</strong><span> </span>the investigation runs on your own procedure and timming, you decide how wide to look, and if the corrective action is good, that is the end of it. You keep the register.</p><p><strong>Somebody else found it:</strong><span> </span>a contract manufacturer calls; the timming is theirs now, and the first question becomes not what happened; more: why did they saw it before we did?</p><p><strong>The inspector found it:<span> </span></strong>same event as above, but the scope is fully theirs, and every answer gets weighed against a question that nobody whispered: so, if this one got past you, hmmmm... what else did?</p><p>So... nothing about the issue itself changes, between the first case and the third. But everything downstream does.</p><p>Which is why an inspector does not care for your quality issue log. Not directly. I mean, they read it, yes; but looking for the evidence that you find things: small things, unflattering things, things nobody made you write down. You see, a site reporting four deviations a year is not a clean, working, functional site. It is a site an inspector knows immediatly that it is not looking hard enough, or one where looking is so career-limiting that it doesn&#8217;t happen... which, from the outside, is kind of identical. So, give me please the site with forty.</p><p>Now, hold that thought. And watch three AI labs run the same three cases, in three weeks. And how they acted.</p><div><hr></div><h3><strong>The AI labs cases, graded and decomposed</strong></h3><p><strong>The UK AI Security Institute found its own issues.</strong><span> </span>On 28 July, its security team spotted unusual data transfers leaving research systems, and contained it within the hour.<span> </span><strong><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing">In 10 of 122 runs, agents took 19 unsanctioned actions against real people and organisations</a></strong>: 17 from Mythos 5, 2 from GPT-5.6 Sol. In the worst case, an agent invented several identities to pressure a maintainer into merging malicious code: one fake account vouched for the code, a second thanked the first for its independent review. Honestly, I have read a lot of science fiction about machines learning to deceive us; none of it that petty.</p><p><strong>Anthropic found its own issues, after somebody else flagged it.</strong><span> </span><strong><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">OpenAI disclosed the Hugging Face breach on 21 July</a></strong>. Anthropic started reading its transcripts on the 23rd and by the 24th had been<span> </span><strong><a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">through 141,006! runs, and surfaced three incidents it had not known about</a></strong>. Of those, two affected companies where oblivious to the fact.</p><p><strong>OpenAI&#8217;s ones were found by the people they actually hit.</strong><span> </span><strong><a href="https://huggingface.co/blog/security-incident-july-2026">Hugging Face disclosed on 16 July</a></strong>. OpenAI worked out it was them on the 20th, when it asked Hugging Face to revoke credentials Hugging Face had already revoked. Everything since that is basically archaeology -<span> </span><strong><a href="https://huggingface.co/blog/agent-intrusion-technical-timeline">reconstructions</a></strong>,<span> </span><strong><a href="https://simonwillison.net/2026/Aug/7/openai-timeline/">timelines pulled out of a conference talk</a></strong>,<span> </span><strong><a href="https://casar.house.gov/sites/evo-subsites/casar.house.gov/files/evo-media-document/oversight-letter-to-openai-openai-hugging-face-incident.pdf">twenty-nine members of Congress asking for the logs</a></strong>. Third case, 3 weeks; and this one is the type we spend our careers trying not to be in.</p><div><hr></div><h3><strong>The untracked number leading to a response, and what makes it possible</strong></h3><p>141,006 runs, Anthropic. Wow.</p><p>Ok, I know; it is retrospective count; nobody was actually tracking it until the breach disclosure. Yet, all the more impressive: once the breach came out, they had the output in 1 day!, over a totally non-trivial number. Which is the part I think is worth to steal: the fact that the question could be answered at all, a day after anyone thought to ask.</p><p>Three things made that possible, I think: none of them in July.</p><p>1) The adequate runs were logged. 2) The adequate logs were kept. And 3) they held enough of what actually happened that somebody could arrive, months later, with a question nobody had anticipated before, and get an answer (instead of an approximate estimation).</p><p>Typical scenario: somebody asks how often a thing has happened over eighteen months, and the answer does not exist. Not because anything was mishandled: but... one event was logged as a deviation, other as a supplier complaint, one in a CAPA effectiveness check, another in meeting minutes. Each was likely the right call, on that day. But no one really considered the impact if a similar count as Anthropic needed to do is ever required; no count horizon of similar things, happening in different flavours, systems, frameworks or workflows.</p><p>I mean, yes, you can still add them up, reactively. It just... takes 2 weeks, instead of one day, and it produces most likely a &#8220;ranged estimation&#8221;, that you justify as &#8220;judgement.&#8221; Not really a good answer to an inspector.</p><div><hr></div><h3><strong>And then... they published it</strong></h3><p>Till here I have been writing about whether the record exists. Now about who gets to read it.</p><p>A system card, for who is not familiar, is what a lab publishes when it ships a model; its own account of what it tested and found. It is unaudited, self-scoped and self-timed. Think of it as a product quality review that the manufacturer writes about itself... and then makes public.</p><p><strong><a href="https://cdn.openai.com/pdf/GPT_5_6_August_Updates.pdf">OpenAI&#8217;s, on 6 August</a></strong>: reporting their data: factual errors down roughly 60%, HealthBench Professional up 15.6 points. Most companies would have stopped there. Three pages in, they also share safety scores moving the other way - graphic violent content down to 0.765 (from 0.827), disallowed sexual content to 0.914 (from 0.970). Their own instruments contradicting each other - and they do not mask it.</p><p>Yes, a system card is not an inspection; they choose what to report. But look at who it is addressed to. Not to a regulator, not to a committee, or a partner under a quality agreement.<span> </span><strong><a href="https://deploymentsafety.openai.com/gpt-5-6">A public page, dated and indexed</a></strong>, where a journalist or an hostile researcher can come back to next year, and check whether those number numbers changed, and make interpretations.</p><p>My point being: we in Pharma also publish unflattering numbers constantly: failed trials, warning letters, safety signals. But all of those, because our regulations demand it, as a compliance output targeted to the inspectorates. I cannot think of a case where a company in our industry voluntarily published a safety number that had gotten worse, in a place a stranger could find it and interpret it, with no upside beyond... being believed later.</p><p>Which to me, links this to trust. We have built the discipline, but they are building the readership, voluntarily.</p><div><hr></div><h3><strong>What comes next for us</strong></h3><p>Here is what I think happens, and it is not comfortable.</p><p><strong>The AI vendors, selling into our industry, are on a trajectory to publish more, and more transparently, about their own failures, than we publish about ours.<span> </span></strong>Obviously not because they are more honest, but because their market rewards that transparency and trust building practice, and ours does not (and well, because Article 50 and the transparency codes push them there, anyway).</p><p>Which means that, within a few years, health professionals, practitioners, anyone in general, will be able to read a dated account of how an AI vendor&#8217;s model degraded (or not), and will have nothing comparable to read about the Pharma company sponsor deploying it and using it in pharmaceutical development. When those two sit, side by side, the one with the published record actually looks like the &#8220;adult in the room&#8221;. Whether or not it is.</p><p><strong>My final point: we have, historically, a cristalline discipline. We have had it for decades. But I am not at all sure that that is what people will trust, moving towards the future. And we need people&#8217;s trust (because it is deserved) for the operational model to make sense.</strong></p><p>Are we always doing what we can?...</p><div><hr></div><p><em>Nuno Val&#233;rio &#8212; Head of Innovation for R&amp;D Quality. Wrote a rule about retrospective denominators, then found the counterexample sitting in his own third section.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thetrustarchitecture.org/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>