Where the System Ends and Judgment Begins
2 September 2026·16 min

Where the System Ends and Judgment Begins

Clinical AI pushes uncertainty to the edge, where someone has to decide whether the system's answer applies. The question becomes whether the organization has decided who may disagree, what that disagreement means, and what happens afterward.

Every evaluation of a clinical AI system describes the ordinary case. Sensitivity, specificity, area under the curve, calibration: these are properties of a distribution, not of any single patient. What they tell us is how a system behaves across a population. They do not, by themselves, tell us how reliably it behaves for the particular patient in front of a clinician, when that patient sits outside the regions of the data the system was actually validated on. The edge is not a smaller version of the center. It is often a different problem altogether, one the system may never have been systematically evaluated against, because the validation data did not adequately represent it. Frequency in the data is not the same thing as importance in practice: the case that occurs most often is rarely the case that matters most when it occurs.

That is worth sitting with, because it is not usually a story about a poorly built model. Validated in aggregate, most of these systems likely perform the way their developers claim. The problem sits specifically at the boundary, in the accumulation of cases the validation set never really represented, and what happens there depends on what kind of boundary case it actually is.

It also helps to be precise about what actually happens at that boundary, because "exception" currently does several jobs at once. Sometimes the model is honestly uncertain, and its confidence tracks the thinness of the underlying data. Sometimes the model is confidently wrong, producing a clean, decisive output for a case its training data never really contained. Sometimes the model is doing exactly what it was built to do, correctly predicting the label it was trained on, and the problem is that the label was never quite the right thing to optimize for in this particular patient. And sometimes there is no fact of the matter to get right, because the clinical situation is genuinely contested, the kind of case where two competent specialists would reasonably disagree with each other regardless of what any model says. These are different problems. Treating all of them as one undifferentiated category, "the exception," is itself part of why institutions struggle to build a coherent response to any of them.

None of this complexity is new. The unusual case, the atypical presentation, the conflicting guideline, was always harder to get right than the routine one; automating the routine case didn't put the difficulty there, and it doesn't remove it. What changes is the appearance of the problem: once the ordinary case is running on its own, there is a natural temptation to assume the system is roughly as reliable everywhere as it is in the region it was actually built and tested for. It isn't. If the edge cases themselves go unexamined on that assumption, whoever eventually stands in front of one is, by default, unexamined too.

The Person Who Receives the Exception

Somewhere in every clinical decision-support deployment, there is a person who receives the case the system did not handle well. Which kind of exception it is, model uncertainty, model error, a mismatched target, or genuine clinical ambiguity, changes what that person actually needs: more information, more time, a second opinion, or simply the standing to say the recommendation does not apply here. Who that person is, and what they are given, is almost never decided with this distinction in mind. In most settings it is simply whoever is on shift when the case arrives.

It is worth being exact about what this moment is. An exception, in most clinical decision-support systems, is not a technical category so much as the point where a person's judgment and the system's output diverge, and someone has to decide which one governs. Exception handling, described this way, is mostly a euphemism. What it names is disagreement, resolved locally, one case at a time, almost always without institutional visibility unless the outcome is bad enough to trigger review.

That reframing does real work. Deciding who receives an exception is, in effect, deciding who is authorized to disagree with the system, and under what conditions. It is a different question from whether an institution trusts the system's output in general. It is a question about whether a particular person's dissent from that output, in a particular case, has been given any standing at all, or whether it will only be recognized after the fact, if the outcome turns out badly enough to warrant a review.

An Override That Exists on Paper

Regulators have started to notice this gap, and it is worth being precise about how they have addressed it, because the precision matters more than the general gesture. Article 14 of the EU AI Act, on human oversight for clinical decision-support systems that fall into its high-risk category, does not just require that a human be nominally in the loop. It requires that whoever is assigned oversight be able to correctly interpret the system's output, remain alert to the tendency to over-rely on it, a risk the Act names explicitly as automation bias, and be able to decide, in a given case, to disregard the output or not use it at all.

That is a genuine legal requirement, not a general right to disagree so much as a mandate that the capability to disagree be actively built into the system rather than left to whatever the interface happens to allow. For a narrow set of high-risk systems elsewhere in the Act, oversight is structural rather than optional: no action may be taken on the system's output until more than one competent person has independently confirmed it. Clinical decision support is not among them. Where a clinical AI system does fall within the Act's high-risk category, the relevant requirement is closer to the first kind of oversight, that a person be capable of disagreeing, not the second, that disagreement be built into how the decision gets made. The practical structure of that disagreement is left to the system and its deployer.

And what fills that space, in practice, is a familiar asymmetry. Accepting a recommendation requires a click. Departing from it usually requires a reason: a free text field, a dropdown of justifications, sometimes a second sign-off. This quietly establishes which action is the default and which is the deviation that needs defending. Compliance is frictionless. Disagreement carries a cost, paid in time and in the small administrative burden of explaining yourself to a system that owes you no explanation in return.

Take the sepsis early-warning tool built into the Epic electronic health record, deployed, in various versions, across hospitals in the United States. It fires an alert when a combination of vital signs and labs crosses a threshold correlated with the eventual diagnosis. One external validation, at two county emergency departments, found a positive predictive value well under ten percent: the overwhelming majority of positive alerts were not, in fact, sepsis. A separate analysis from University of Michigan Health went further, and found that part of what the model was detecting was not the underlying physiology of sepsis at all, but the fact that a clinician already suspected it. Some of the data it relied on, a blood culture ordered, antibiotics already started, existed only because someone had already begun treating the patient as though they had the condition.

This is where the cost of disagreement and the routineness of disagreement turn out not to contradict each other, but to describe two ends of the same undifferentiated interface. Where an alert is rare and the stakes feel unusually high, the added friction of justifying an override is enough to make compliance the path of least resistance. Where an alert is frequent and mostly wrong, as with the sepsis tool, the friction stops mattering well before the fatigue does: a clinician only has to click through a handful of false alarms before all of them start to look the same, and disagreement stops being a decision and becomes a reflex.

Published studies of clinical decision support report override rates for medication and dosing alerts ranging from under half to well over ninety percent, and only a minority of overrides are ever formally assessed for whether they were clinically appropriate. Either way, the interface fails to track what actually distinguishes one exception from another, and either way, what should be a deliberate act of judgment collapses into something else.

Override is not evidence the clinician was right. It is evidence that the system and the clinician disagreed, and that disagreement is itself a signal worth examining, whether it turns out to reflect a genuine model failure, a workflow problem, or simple fatigue. The reverse holds too, and matters more than it first appears. Compliance is not evidence the system was right either. Given what accepting a recommendation costs, nothing, and what departing from it costs, time, justification, exposure, a clinician may simply accept a recommendation because contesting it wasn't worth the trouble, not because it was the better of the available options. The asymmetry built into the interface doesn't just suppress disagreement. It manufactures agreement that looks, in the data, identical to genuine concurrence. What makes any given case informative is not whether the clinician complied or overrode, but what kind of divergence, or non-divergence, actually produced that outcome, and that is precisely the distinction most interfaces are never built to make.

Legitimacy, Not Ability

The ability to override a system and the institutional legitimacy to do so are not the same thing, and treating them as equivalent is where most oversight regimes go quietly wrong. A button that lets a clinician depart from a recommendation says nothing about what happens to them afterward if they do, and that second question, mostly invisible to any technical audit, is the one that actually governs behavior.

The anthropologist Madeleine Elish has a useful phrase for what tends to happen when something goes wrong at the boundary between a person and an automated system: a moral crumple zone. Just as a car's crumple zone is engineered to absorb the physical force of a collision so the passenger doesn't have to, the human operator in a complex automated system can absorb the moral and legal force of a failure that was never fully within their control. When a clinician follows a recommendation and the outcome is bad, responsibility tends to diffuse into the system, the protocol, the standard of care at the time. When a clinician overrides a recommendation and the outcome is bad, responsibility concentrates, sharply, on the person who chose to depart. That does not mean the clinician should be shielded from the consequences of their own decisions. It means responsibility should track the actual distribution of control across the system and the people using it, rather than defaulting to whoever happened to make the final click.

That asymmetry does not need to be written into any policy to shape behavior. It only needs to be understood by the people making the decision, and clinicians understand it well. The safest action, in a system built this way, is agreement, not because agreement is usually right, but because it is the only action that does not leave the person personally exposed if it turns out to be wrong.

There is an inversion worth noticing here. The feature that a regulator or an auditor reads as proof of meaningful human control, the visible override option, the documented justification field, can be, from the other side of the same screen, exactly what makes exercising that control expensive enough to avoid. The artifact serves two audiences and tells them opposite things: to the institution, evidence that oversight exists; to the clinician exercising it, a reason to think twice before doing so.

Underneath this is a simpler way to describe what these systems actually assign to people. Every automated system leaves some uncertainty unresolved. The question a workflow design answers, whether anyone means to answer it or not, is who absorbs that uncertainty. When a system handles the ordinary case and a person handles what is left over, that person is not merely "in the loop." They are the organization's residual uncertainty processor, and how much authority, protection, or exposure comes with that role is a design choice like any other, even when nobody has consciously made it.

What the Exception Reveals

Framed that way, an institution facing this residual uncertainty has real options, and most of them are already visible in how a given system is built. It can give the person who absorbs the uncertainty real authority and the resources to use it. It can make their disagreement procedurally expensive, through documentation burden or workflow friction. It can make them personally liable for outcomes they only partly controlled. Or it can build a way to learn, systematically, from what they encounter. These are not mutually exclusive, but many institutions, without quite deciding to, default toward the middle two and neglect the fourth.

Learning from the fourth option requires more than logging the override, which is the part almost every system already does, because logging is cheap and defensible. It requires treating different patterns of override as different diagnoses. An override that happens rarely, for a specific and unusual patient, suggests a genuine exception worth studying on its own terms. An override that happens constantly suggests a poorly calibrated model or an unworkable workflow, not a string of individually meaningful judgment calls. An override clustered around one clinician suggests a training or trust problem. One clustered around a particular patient subgroup suggests a blind spot in the model itself. Collapsed into a single undifferentiated category, "override," none of these patterns is visible, and the log becomes, by default, a record kept in case someone asks rather than a source of anything to learn.

Go back to the sepsis tool. Whether the hospitals running it treat their override logs as an audit trail or as a diagnostic instrument determines whether anyone ever connects the dismissal rate to the specific finding that the model was partly reading clinician suspicion back to itself. Both are visible in the same data. Only one requires anyone to look. Vendors of these systems tend to sell them on two claims at once: that they are already reliable, and that they will keep improving as more real-world data comes in. The second claim depends entirely on the fourth option, some structured path from a clinician's disagreement back into how the system or the workflow around it gets revised. In practice, that path is usually the first thing missing. A system can run continuously without learning continuously; the log exists, the analysis of it typically does not.

A system is not finished when it produces a reliable answer for the ordinary case. It is finished, if that word ever really applies, when the organization has decided what happens once someone competent believes the answer is wrong.

Many institutions haven't made that decision explicitly. They have simply let it get made by default, by whoever happens to be holding the pager, by what a justification field costs to fill in, by what happens to the clinician who disagrees and turns out to be wrong. None of that is a failure of the model. It is what happens when an institution treats its own edge cases as a nuisance to be logged, rather than as the one place it could actually find out whether the system deserves the trust it has already been given.

Related

1 / 3