What does it actually take to ship AI in a regulated industry?
Provenance on every component, not a single passing benchmark. Corti CEO Andreas Cleve's requirement is that you can explain what each part of the pipeline did, why it changed, and what went into it. He also argues that public leaderboards are close to useless as evidence, because the models topping medical benchmarks are the ones nobody ships.
Enterprise AIAI Evaluation & Reliability
1. The slowness is a bar, not a blind spot
Andreas Cleve, CEO of Corti, rejects the premise that regulated industries are behind. On healthcare he is blunt that the laggard story is “super boring,” and offers the counter-observation that a clinic packs more devices per room than his own machine-learning office does. His reading: “they adopt tons of technology, and they do it quite fast, but they have a pretty high bar for doing it.”
That reframes the problem. The obstacle is not appetite, it is the evidence you have to produce before anyone will let the thing touch a patient.
2. Every component needs its own benchmark
The concrete requirement is provenance across the pipeline rather than a score at the end. Cleve describes it as a factory: “In the factory, every component that’s interacting has to have benchmarks, so we need to understand remove data, what happens, where does it happen?”
And the artifact regulators and buyers want is an explanation, not a number: “we need to be able to actually explain what we do, why we did it, why we changed it, when we did, what’s in the model cards, like what’s in the food we’re serving, what’s the ingredients we used, why do we use them, when do we move them, how do we launch it, where is it hosted, at what level and what state are we encrypting.”
His summary of the shape of the work is worth sitting with if you are scoping a regulated build: “you can’t win health care by being good at one thing. You have to be good at many things.”
3. Public benchmarks are weak evidence, by his own argument
Cleve makes an observation about medical leaderboards that cuts against the usual proof point. On the big public health datasets, “the leaders are noncommercial models.” His inference is hedged and he says so: “It’s probably because it’s not very useful in practice… there’s probably really good reasons why it was trained in a way that makes it really good at tests but not very good in the wild.”
What replaces the leaderboard is layered checking. He describes “a series of models as judges passing judgment on whether or not this reasoning pipeline gets it” rather than batch processing, plus a large stock of holdouts. See what is LLM-as-a-judge for when that check can be trusted, and how do you evaluate an AI system when the output isn’t deterministic for the surrounding method.
4. The workflow can hand you a label for free
The part worth copying: Corti’s product sits in a workflow where a clinician has to approve the output. That approval is a graded signal, not just a UI step. “a doctor has to say, what you did was cool. So we actually at the end get a quite cool golden label of whether or not it’s as complete or as concise as somebody would actually use it for.”
If a regulated workflow already requires human sign-off, that gate is also an evaluation set nobody had to build.
5. Compliance turned out to sell
Cleve reports the inversion directly: “Being from Europe, we’re really good at compliance… If I were doing AI for many other things in the world, I wouldn’t be saying this. It would be an inhibitor.” Instead a large US customer with a long purchasing cycle came to them, and part of what they liked was alignment with the EU AI Act and GDPR, on the buyer’s expectation that “no matter what the market will drift towards some of this stuff.”
He does not oversell it. He grants that some of the red tape “we should remove,” while holding that under it sits “a nucleus of truth” about needing to trust the system in a high-stress environment.
6. Time saved is not the same as work removed
The counterweight in the episode is about AI scribes. Corti surveyed 2,000 early adopters in the US, and Cleve reports “they told us on average they spend at least three hours a week just, like, vetting and correcting their AI,” alongside a standing worry that something was misrecorded. His verdict on that pattern is that “we’re creating workflows here that shouldn’t be there.”
A tool that shifts effort from writing to checking has not necessarily paid off yet. That is the same measurement trap covered in how do you measure whether AI is actually paying off.
Why it matters
Regulated buyers are not asking whether your model is good. They are asking you to show your work per component, and to keep showing it after every change. Teams that treat compliance as a launch gate discover it is a build constraint; teams that build for it early, on Cleve’s account, occasionally find it closing deals instead. Where the data itself cannot leave the building, on-premise deployment is the adjacent question.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.