Somewhere, right now, a pharmacovigilance system is recalculating a drug's benefit-risk ratio. A new batch of adverse event reports has come in; the model compares the observed rate against what background reporting would predict, updates its estimate of how the risk side of the ratio has shifted, and returns a new number. This takes about as long as the query needs to load. What happens next takes months, sometimes years.
The updated ratio has to be reviewed by a safety scientist, then presented to a causality assessment committee that will apply criteria established years, sometimes decades, earlier by people no longer in the room to explain why, and whose current members will disagree with each other along the way, sometimes for good reasons and sometimes out of professional caution. If the committee agrees the new number changes anything, the finding travels to a regulator, who opens its own review, invites the manufacturer to respond, and eventually decides whether the ratio still clears the bar for the drug to stay on the market unchanged, or whether the label, the population it's approved for, or its availability needs to shift. Somewhere between the model returning its new number and anything actually changing, a year has usually passed.
It is tempting to experience this gap as bureaucratic drag, the kind of friction a sufficiently motivated reformer could engineer away. But what if the gap is not a design flaw. What if it is the point of the system. The model's job was to produce a number. The committee's job was never to reproduce that number faster. Its job was to decide whether that number, on its own, was enough to change what a regulator, a manufacturer, a prescriber, and a patient are all willing to accept, and that is a different kind of work entirely.
Underneath that gap sits a specific order of operations. Several questions, each one only visible once the one in front of it has been answered.
The Manufacture of Acceptance
Start with the first question. A recalculated benefit-risk ratio, however well the model behind it is validated, is not itself a decision. It is an input that becomes a decision only once enough people accept it as grounds for one, and manufacturing that acceptance is a much older institutional function than pharmacovigilance itself. Courts, regulators, scientific peer review, ethics committees, clinical boards: none of them exist primarily to maximize intelligence. They exist to manufacture legitimacy, the quality that lets a decision stick even among people who disagree with it. Put plainly: institutions were never primarily intelligence engines. They are legitimacy engines, which is why so much of the current argument about whether machines can replace them misses the point; the machines were never competing on the axis that mattered.
A judge's verdict and a confession extracted under duress might occasionally point to the same factual conclusion. Only one of them is a decision anyone is obligated to respect. Correctness and acceptance have never been the same currency, but for most of institutional history we could treat them as roughly interchangeable, because producing a defensible answer was itself difficult enough that whoever did the work earned a claim to legitimacy in the process. Expertise was scarce, so the expert's authority and the expert's correctness arrived bundled together.
That bundling is what is coming apart. AI has made a specific kind of work radically cheaper: producing a plausible, often well-calibrated number or draft, a recalculated benefit-risk ratio, a diagnostic likelihood, an actuarial price, a legal risk memo, across an increasing range of domains. None of this required scarce expertise to produce a defensible first answer, it required a sufficiently large model and a reasonably well-posed question. What has not gotten any cheaper is the machinery that turns that number into an accepted decision: representation, the chance to object, deliberation among people who bear the consequences, someone who can be named if the decision turns out wrong.
This results in a sometimes huge distance between the moment a system is technically capable of producing a number and the moment an institution is prepared to stand behind it. In healthcare AI specifically, this shows up as the gap between deployment and embedding. A model recalculating a benefit-risk ratio in real time can clear its validation study, get built into a production pharmacovigilance pipeline, and sit there, technically deployed, for a long stretch before anyone routes an actual regulatory decision through its output, before it is embedded into how safety review actually happens. The validation study answers whether the new ratio is accurate. It does not, and cannot, answer whether the people whose licenses are on the line are willing to treat that ratio as grounds for changing what stays on the market. No acceleration of the underlying model closes that gap, because the gap was never about the model's intelligence.
You can see the same gap between knowing and accepting well outside regulated medicine. Nutrition science has known for decades, with about as much confidence as population-level science offers, that high sugar intake tracks with obesity, type 2 diabetes, and cardiovascular disease. Sugar taxes remain contested nearly everywhere they are proposed anyway, not because anyone seriously disputes the underlying biology, but because turning that biology into policy requires legitimacy among consumers, manufacturers, and legislators that the biology alone cannot supply. Pharmacovigilance committees exist for exactly the same reason, scaled down to a single drug and a single number: not to recompute a ratio the model has already recomputed, but to convert that number into a decision enough people will regard as binding.
An obvious objection follows. Couldn't AI eventually manufacture legitimacy too, synthetic public comment analysis, model-run consultations, automated appeals review? It could automate legitimacy's paperwork. It could not automate who is on the hook when the decision turns out wrong, and a judge, a committee member, or a regulator can be appealed, recused, or sued in a way no model can. Legitimacy is inseparable from that exposure, which is exactly what stays expensive once everything else gets cheap.
Paradoxically, making the ratio cheaper to compute does not reduce the demand for governance around it. It increases it. Institutions now have to process a growing volume of increasingly plausible new numbers while still preserving accountability for whatever decision follows each one, and preserving accountability at higher throughput is harder, not easier, than preserving it at the old, slower pace.
What the Ratio Is For
Suppose the committee accepts the new ratio: the legitimacy problem is solved, for this case. The second question was hiding underneath the first one the whole time. Accept it relative to what? A benefit-risk ratio only means something once someone has decided what counts as benefit and what counts as risk, and how much of one is worth trading for how much of the other: reducing adverse events, preserving access for patients with no alternative, protecting the incentives that fund future research. Change the weighting and the same recalculated number can point toward opposite decisions.
AI is extraordinarily good at optimizing within a stated goal. Give it an objective function and it will typically find better solutions faster than a person would by hand, whether the goal is winning at chess, folding a protein, minimizing default risk, or maximizing the sensitivity of a diagnostic test. What it is comparatively bad at is telling you that you specified the wrong objective in the first place, because from inside an optimization process, the objective is not a hypothesis. It is a given, and a very good optimizer will pursue a bad objective just as relentlessly as a good one. A model can tell you precisely where a given weighting of the ratio lands. It cannot tell you that you weighted it wrong.
Cancer screening supplies an uncomfortable illustration of the same failure elsewhere in medicine. Several national screening programs spent years optimizing for one clear target: detect more cancer, earlier. Sensitivity climbed, detection rates climbed, and a meaningful share of what got caught, some low-grade thyroid and prostate findings among them, would never have caused harm within the patient's lifetime. Nobody had gone back to ask whether "detect more" was still the objective that mattered once detection got cheap, as opposed to something closer to "extend meaningful life." A benefit-risk ratio can drift the same way. A threshold built around minimizing missed safety signals will look different from one built around minimizing unnecessary withdrawals, and nothing in the model that computes the ratio tells you which one you actually meant.
Pharmacovigilance makes the underlying problem concrete. "Reduce adverse events" sounds like a goal, but it is not one on its own; a ratio single-mindedly optimized for that would eventually justify withdrawing almost anything, since every functioning drug carries some risk. The real work of a safety committee is deciding how much risk is acceptable in exchange for how much benefit, for which patients, compared to what alternatives, and that weighting is exactly the kind of choice a model cannot make on the committee's behalf, because the model was only ever asked to compute the ratio, not to decide what the ratio was for.
This is where the first two problems meet. Deciding what a benefit-risk ratio should be optimizing is not a technical act. It is a legitimacy-laden one, because whoever gets to decide how much risk is acceptable, for which patients, is exercising precisely the kind of authority that regulators and boards exist to allocate. As answers get cheaper, the contest does not disappear. It moves up a level, from is this the right number, to who gets to decide what the number should be for, a question that, as it turns out, institutions can quietly lose the answer to even after they have made it.
The Threshold No One Remembers
Push the regress one step further. Suppose a committee has both the legitimacy to act and a clear, deliberately chosen answer to what the ratio should optimize for. A third problem appears, quieter than the first two but perhaps more consequential: over time, institutions lose track of why that weighting was chosen in the first place.
There's a well-known joke about an answer to the ultimate question of everything turning out to be a meaningless number, because nobody remembers the actual question. Something like it plays out, less amusingly, inside organizations whose institutional memory does not survive a leadership change, a merger, or a decade. They keep the answer: a threshold, a screening protocol, a model quietly running in production. They lose the reasoning that justified it. Ask why a given benefit-risk threshold sits exactly where it does, and you will often get an answer about precedent rather than purpose, the number a committee agreed on in some earlier year, for reasons nobody currently at the table was present for.
Governments that manage long-lived nuclear waste have run into an extreme version of the same problem. Some sites need to stay marked as dangerous for longer than any currently spoken language, symbol system, or institution has existed, which means no ordinary warning sign can be trusted to survive the span it needs to cover. The proposals that came out of thinking seriously about this, elaborate monuments, deliberately unsettling landscapes, a designated caste of future custodians tasked with keeping the memory alive across generations, sound eccentric until you notice they are solving exactly the problem this essay is about: an answer, keep away, that is worthless unless something also survives to carry the reason why.
The same shift is visible at civilizational scale. Historically, the limiting question was almost always capability: can we build this. That scarcity did a great deal of our deciding for us, because for most of civilization's history we simply couldn't. Building has since gotten dramatically cheaper, and the limiting question has visibly shifted from can we build this to should we build this, with a third question already behind it: who gets to decide, and will the reasoning still be legible to whoever inherits it.
This, in the end, is the same problem as a pharmacovigilance committee inheriting a benefit-risk threshold nobody currently in the room helped set. The model's new number arrives instantly. The organization's capacity to say why the old threshold sits where it does did not arrive with it, and if that capacity has quietly eroded in the years between, no amount of additional intelligence, artificial or otherwise, restores it. Faster numbers do not, by themselves, preserve the reasons a threshold was drawn where it was.
Upstream
None of this is an argument against building capable, fast, cheap systems. Speed is not the problem. The mistake would be assuming that a cheaper number shrinks the amount of expensive human work required, when what it actually does is relocate it. From computing the ratio, to accepting it. From accepting it, to deciding what it should be optimizing for in the first place. From deciding, to remembering, months or years later, why the threshold was drawn where it was.
Institutions increasingly look inefficient only because we keep judging them by how quickly they produce an answer. That was their bottleneck for most of history. It is a shrinking one now. What is left is exactly the climb this essay has been describing: deciding whether a number deserves acceptance, choosing what it should optimize, and preserving the memory of why, long after the people who chose it have left the room. Those functions are not made obsolete by cheap intelligence. They become more valuable, not less, the more abundant intelligence gets.
Legitimacy, judgment about objectives, and institutional memory were never a computation problem, which is why cheapening intelligence did not cheapen them. The machines are producing new numbers faster than we can agree on what to do with them. That gap, not any shortage of intelligence, is the scarcity upstream of it, and it is where the real work now lives.

