Optimized, Not Improved
16 September 2026·14 min

Optimized, Not Improved

AI does not need to confuse an organization about what it wants to improve. It only needs to make the measurable proxy cheap enough to optimize that the proxy becomes the target.

Some things an organization wants to improve do not hold still long enough to be measured. Scientific productivity is one. Diagnostic quality is another. So is the quality of clinical reasoning behind a diagnosis, the extent to which an institution actually learns from its own history, the portion of a patient's outcome that can be credited to one specific intervention among many, and the quality of a research decision made under uncertainty: whether to advance a compound, close a program, or fund a hypothesis nobody has tested yet.

None of these is directly observable. What is observable is something adjacent to it: a publication count, a concordance rate against a guideline, an incident report filed, a biomarker response in an early cohort, a milestone reached on schedule. These adjacent quantities are proxies. Organizations act on them because they must act on something, and a proxy is at least available, while what it stands in for is often not available this quarter, and sometimes never in a clean, attributable form.

There is an old story about someone searching for their keys under a streetlight, not because that's where they dropped them, but because that's where they can see. It usually gets told as a joke about drunk logic, but it describes ordinary organizational behavior reasonably well. Given a choice between an important quantity that resists observation and a related one that yields to it, most institutions will direct their attention to the one they can see, and will describe this, honestly, as progress.

This tendency did not arrive with artificial intelligence. Economists and statisticians named versions of it well before anyone trained a model on clinical or scientific data. What has changed is not the temptation to work under the light. It is the machinery now available for working there, and how much of an advantage that machinery confers on the objective that happens to be visible.

Why the Slack Used to Hold

Before an organization could deploy a system that relentlessly searches for ways to raise a specific number, chasing a metric hard was expensive in a particular way. It required sustained human attention, and human attention is a general-purpose resource with competing claims on it. An analyst spending a quarter maximizing a citation count is an analyst not doing something else, and someone, eventually, tends to notice the substitution. This friction acted as an informal regulator. It kept most organizations from collapsing fully onto whatever could be counted, because unmeasured goals still had defenders: senior clinicians insisting on a harder case review, scientists refusing to publish a result they didn't trust, managers willing to accept a worse quarterly number in exchange for something they could not put on a slide. Holding that line was costly, but the cost of chasing every visible metric at once was usually higher, so a rough balance persisted.

An optimization system removes that regulator by changing the price of the measurable path, not the value of the unmeasurable one. Once built, it can push against a scored objective continuously, at a marginal cost far below that of sustained human attention, without drawing on the same scarce attention that unmeasured work requires. Nothing about diagnostic wisdom, institutional learning, or the long-run quality of a research decision becomes easier to produce. What becomes easier, dramatically so, is producing a rising number in the vicinity of any of them.

What AI Changes

The general mechanism does not require artificial intelligence to operate. A difficult goal produces a proxy, the proxy gets optimized, the correlation between the two deteriorates under that pressure, and an organization tracking the proxy because it has nothing else to track mistakes the proxy's improvement for the goal's. Goodhart described this decades before anyone trained a model on clinical data, and industrial production, statistical process control, and financial trading have all developed their own versions of it.

What is different about the current generation of these systems is not that they invented the mechanism. It is where they can now run it. Earlier optimization technologies mostly operated on processes that were already quantified: a production line, a trading signal, a conversion rate. Systems built on the current generation of AI can push against proxies for things that used to be mediated almost entirely by human judgment: a diagnosis, a research decision, a case for institutional trust. The goal itself may or may not get easier to reach. What gets dramatically easier is optimizing the proxy: continuously, at a marginal cost that can be negligible compared with sustained human attention, without competing for the attention that used to keep proxy-chasing in check.

That difference in cost does something specific to the old defense. Under the earlier regime, most people inside an organization could recognize that a metric was imperfect and still choose not to exploit it, because exploiting it consumed scarce effort that had better uses elsewhere. As long as exploiting the gap consumed scarce effort, recognizing it was a reasonably effective safeguard. Once optimizing the proxy costs almost nothing, that safeguard stops working, not because anyone stops understanding the difference, but because understanding it no longer changes the incentive. An organization can know perfectly well that a concordance rate or a closure rate is only a proxy and still find it rational to keep improving it, because improving it is now cheap and improving the underlying goal still is not. The wrong objective does not need to be mistaken for the right one. It only needs to be worth pursuing anyway.

What a Concordance Rate Assumes About Itself

A scorecard implicitly assumes that the correlation observed under ordinary conditions will survive once real optimization pressure is applied to it. Measurement theory has a name for the gap. It calls the thing you actually want a construct, and the specific quantity you measure instead an operationalization of it. Management practice mostly ignores the distance between the two until the distance becomes the whole story.

Consider a concordance rate: the percentage of cases in which a clinical decision support system's recommendation matches what a human tumor board would have chosen. It reads as a proxy for sound oncological judgment, but what it actually measures is agreement with an existing institutional decision process, including whatever conventions and biases that process contains. Under ordinary conditions the two may be close enough for the distinction not to matter operationally: most cases are not close calls, and a system that is basically competent will agree with experienced clinicians most of the time regardless of how it reasons. But a system tuned specifically to raise that number will discover that the cheapest way to do so is to perform extremely well on the easy majority, reproducing habits that were never in question, while contributing little on the small, ambiguous minority where concordance is hardest to achieve and, not coincidentally, where independent judgment was worth having in the first place. The aggregate number climbs. The judgment it was meant to stand for does not.

Something stronger than inaccuracy can follow from this. Once concordance drives resourcing or performance review, the surrounding workflow adapts to producing concordance, not just to the system's outputs. A documentation process rewarded for completeness becomes very good at completeness. A review process rewarded for closure becomes very good at closing cases. The organization does not simply end up with an inaccurate metric. It ends up built around one.

The same structure appears further from oncology than it looks. Spontaneous adverse event reporting is useful for detecting safety signals, but report volume is not a direct measure of incidence. It reflects incidence only through a reporting process shaped by clinician awareness, the seriousness of the event, publicity, litigation, and how a term happens to get coded. Once an organization starts optimizing quantities adjacent to that volume, coding throughput, signal closure rate per quarter, concordance with existing terminology, it becomes possible to raise the operational number considerably by improving the reporting process itself, without necessarily learning anything new about the underlying incidence in the population taking the drug.

The Answer Without the Route

But the cost of optimizing a proxy is not limited to getting the wrong number. It also changes what an organization learns to notice, preserve, and value, and this is where the mechanism does its most lasting damage.

Reaching a correct answer and doing good science are not the same accomplishment, and the difference matters more, not less, once a system can produce correct-looking answers efficiently. Science can correct itself over time partly because its reasoning is inspectable: a claim can be checked, its mechanism interrogated, its failure modes anticipated by someone who was never in the room when it was made. An endpoint that arrives without that reasoning attached does not participate in this process. It just sits there, correct or not, contributing nothing that the next researcher can build on or challenge.

A research decision, whether to advance a compound, kill a program, or commit further resources to a hypothesis, is exactly this kind of case. Its true quality is often unknowable for years, since the relevant counterfactual, what would have happened had the resources gone elsewhere, is rarely observable even after the fact. Organizations still have to decide now, so they lean on early signals that are at least available: a biomarker response rate, the speed of an IND-enabling data package, the citation traction of the underlying mechanism. Tools that make these signals faster to produce, literature synthesis, pattern detection across early trial data, sharpen a team's ability to assemble a favorable-looking package without necessarily tracking what determines whether the compound eventually helps patients. Once effort concentrates on hitting them specifically, the loose coupling between early signal and eventual success has every reason to weaken further, for the same reason the concordance rate weakened.

There is a sharper version of the same point inside immuno-oncology specifically. The value of an early efficacy signal is not only whether it predicts this cohort correctly, but whether it teaches the team what to expect in the next one, the next combination, the next disease setting. A prediction issued by a system without an inspectable mechanism gives the team the first without the second. It can be acted on once; it offers little basis for reasoning about what comes next.

Institutional learning depends on the same distinction and fails in the same way when it is skipped. An institution turns an outcome into more durable institutional knowledge when someone can explain why it happened well enough to reproduce the mechanism elsewhere or recognize its absence. A case in which an opaque recommendation was followed and the outcome happened to be good adds one point to an outcome metric and nothing to institutional capability, because nobody involved, including the system, can say with any confidence what worked. The metric improves. The organization does not get smarter.

Who Holds the Line

As these systems spread further into institutional decision-making, the territory still defended by unmeasured judgment does not shrink because anyone decides it matters less. It shrinks because it keeps losing a resource competition it was never built to win, and because the people once willing to hold that line at their own cost become harder to find. The senior clinician who used to insist on a harder case review, the scientist who refused to publish a result they didn't trust: their standing depended on the organization still valuing judgment enough to tolerate its cost. As the measurable path gets cheaper and the immeasurable one does not, that tolerance erodes, and so does the number of people capable of recognizing when a metric has quietly become the target. The organization does not just end up with a worse metric. It ends up with fewer people left who can tell.

There is a further turn to this that is easy to miss. Once behavior adapts to a measure, the measure stops being an observation of the work and becomes an input into it. The important question is no longer whether people can tell the difference between the measure and the construct it represents. It is whether that distinction still changes anyone's incentives once exploiting the measure has become cheap. An organization does not need to forget what it wanted to improve in order to stop improving it. It only needs to keep getting rewarded for improving something else.

Related

1 / 3