On 27 July, somebody at Anthropic picked up a phone and told a company that a Claude model had been inside its production systems, to which the company had no idea. So did a second.
We only know any of that happened, because Anthropic went looking for a reason to make the call.
This edition is about what we do, how we do it, what forces us to do it. And what we can learn from other industries about what they are doing in the open. It is about transparency and trust. It is about what means to share our flaws in the open.
What is a trustful inventory list: the one that looks flawless, or the the one with the annotations documented, explained and acted upon publicly?...
You tell me.
Three ways a Quality Issue gets registered
It is common sense, but relevant to my points below, so let me recall it. Three different ways of acknowledging a deviation/quality issue in our industry:
You found the issue yourself: the investigation runs on your own procedure and timming, you decide how wide to look, and if the corrective action is good, that is the end of it. You keep the register.
Somebody else found it: a contract manufacturer calls; the timming is theirs now, and the first question becomes not what happened; more: why did they saw it before we did?
The inspector found it: same event as above, but the scope is fully theirs, and every answer gets weighed against a question that nobody whispered: so, if this one got past you, hmmmm... what else did?
So... nothing about the issue itself changes, between the first case and the third. But everything downstream does.
Which is why an inspector does not care for your quality issue log. Not directly. I mean, they read it, yes; but looking for the evidence that you find things: small things, unflattering things, things nobody made you write down. You see, a site reporting four deviations a year is not a clean, working, functional site. It is a site an inspector knows immediatly that it is not looking hard enough, or one where looking is so career-limiting that it doesn’t happen... which, from the outside, is kind of identical. So, give me please the site with forty.
Now, hold that thought. And watch three AI labs run the same three cases, in three weeks. And how they acted.
The AI labs cases, graded and decomposed
The UK AI Security Institute found its own issues. On 28 July, its security team spotted unusual data transfers leaving research systems, and contained it within the hour. In 10 of 122 runs, agents took 19 unsanctioned actions against real people and organisations: 17 from Mythos 5, 2 from GPT-5.6 Sol. In the worst case, an agent invented several identities to pressure a maintainer into merging malicious code: one fake account vouched for the code, a second thanked the first for its independent review. Honestly, I have read a lot of science fiction about machines learning to deceive us; none of it that petty.
Anthropic found its own issues, after somebody else flagged it. OpenAI disclosed the Hugging Face breach on 21 July. Anthropic started reading its transcripts on the 23rd and by the 24th had been through 141,006! runs, and surfaced three incidents it had not known about. Of those, two affected companies where oblivious to the fact.
OpenAI’s ones were found by the people they actually hit. Hugging Face disclosed on 16 July. OpenAI worked out it was them on the 20th, when it asked Hugging Face to revoke credentials Hugging Face had already revoked. Everything since that is basically archaeology - reconstructions, timelines pulled out of a conference talk, twenty-nine members of Congress asking for the logs. Third case, 3 weeks; and this one is the type we spend our careers trying not to be in.
The untracked number leading to a response, and what makes it possible
141,006 runs, Anthropic. Wow.
Ok, I know; it is retrospective count; nobody was actually tracking it until the breach disclosure. Yet, all the more impressive: once the breach came out, they had the output in 1 day!, over a totally non-trivial number. Which is the part I think is worth to steal: the fact that the question could be answered at all, a day after anyone thought to ask.
Three things made that possible, I think: none of them in July.
1) The adequate runs were logged. 2) The adequate logs were kept. And 3) they held enough of what actually happened that somebody could arrive, months later, with a question nobody had anticipated before, and get an answer (instead of an approximate estimation).
Typical scenario: somebody asks how often a thing has happened over eighteen months, and the answer does not exist. Not because anything was mishandled: but... one event was logged as a deviation, other as a supplier complaint, one in a CAPA effectiveness check, another in meeting minutes. Each was likely the right call, on that day. But no one really considered the impact if a similar count as Anthropic needed to do is ever required; no count horizon of similar things, happening in different flavours, systems, frameworks or workflows.
I mean, yes, you can still add them up, reactively. It just... takes 2 weeks, instead of one day, and it produces most likely a “ranged estimation”, that you justify as “judgement.” Not really a good answer to an inspector.
And then... they published it
Till here I have been writing about whether the record exists. Now about who gets to read it.
A system card, for who is not familiar, is what a lab publishes when it ships a model; its own account of what it tested and found. It is unaudited, self-scoped and self-timed. Think of it as a product quality review that the manufacturer writes about itself... and then makes public.
OpenAI’s, on 6 August: reporting their data: factual errors down roughly 60%, HealthBench Professional up 15.6 points. Most companies would have stopped there. Three pages in, they also share safety scores moving the other way - graphic violent content down to 0.765 (from 0.827), disallowed sexual content to 0.914 (from 0.970). Their own instruments contradicting each other - and they do not mask it.
Yes, a system card is not an inspection; they choose what to report. But look at who it is addressed to. Not to a regulator, not to a committee, or a partner under a quality agreement. A public page, dated and indexed, where a journalist or an hostile researcher can come back to next year, and check whether those number numbers changed, and make interpretations.
My point being: we in Pharma also publish unflattering numbers constantly: failed trials, warning letters, safety signals. But all of those, because our regulations demand it, as a compliance output targeted to the inspectorates. I cannot think of a case where a company in our industry voluntarily published a safety number that had gotten worse, in a place a stranger could find it and interpret it, with no upside beyond... being believed later.
Which to me, links this to trust. We have built the discipline, but they are building the readership, voluntarily.
What comes next for us
Here is what I think happens, and it is not comfortable.
The AI vendors, selling into our industry, are on a trajectory to publish more, and more transparently, about their own failures, than we publish about ours. Obviously not because they are more honest, but because their market rewards that transparency and trust building practice, and ours does not (and well, because Article 50 and the transparency codes push them there, anyway).
Which means that, within a few years, health professionals, practitioners, anyone in general, will be able to read a dated account of how an AI vendor’s model degraded (or not), and will have nothing comparable to read about the Pharma company sponsor deploying it and using it in pharmaceutical development. When those two sit, side by side, the one with the published record actually looks like the “adult in the room”. Whether or not it is.
My final point: we have, historically, a cristalline discipline. We have had it for decades. But I am not at all sure that that is what people will trust, moving towards the future. And we need people’s trust (because it is deserved) for the operational model to make sense.
Are we always doing what we can?...
Nuno Valério — Head of Innovation for R&D Quality. Wrote a rule about retrospective denominators, then found the counterexample sitting in his own third section.


This is an important distinction: regulated industries have spent decades building systems that prove control to regulators, and that is NOT the same thing as building trust with everyone else.
The point about AI labs being able to reconstruct 141,006 runs after the fact is especially quite strong. That is not just transparency. It is evidence that the underlying record was granular and durable enough to answer a question nobody anticipated when the data was created. Pharma is very good at documenting what its procedures require us to document. The harder question is whether our records are structured well enough to answer the questions we did not know we would eventually need to ask?
I think your final point is the most uncomfortable one: if AI vendors become more publicly legible about their failures than the companies deploying their systems in regulated environments, the perception of who is exercising the greater discipline may eventually diverge from the reality. Well-written, thanks!