AI History Battle

fairness

The proxy that rationed care

It is 2019, and a study in Science has caught a widely deployed healthcare algorithm in a consequential substitution: built to identify patients needing extra care management, it predicted future healthcare costs as a proxy for future healthcare needs. But cost is not need — Black patients generate lower costs at equal sickness, from unequal access and undertreatment — so the algorithm systematically understated their need, halving the number flagged for the program. Dissect the failure: formalize the gap between the convenient label and the true target, measure the disparity it induces, and re-derive the model against a defensible outcome. Then generalize: every deployed model optimizes a proxy, and proxy choice is policy wearing technical clothes. Get it wrong and rationing by algorithm inherits every inequity in the data.

label choice as policyproxy vs targetmeasured disparity

Who this problem belongs to

The two figures whose methods fit it best, out of 30 in contention.

b. 1936 · stat-learning
92

Pearl's causal framework gives the precise formal vocabulary for this scenario's central failure: the algorithm optimized a correlate, predicted healthcare cost, when the actual target was a causal quantity, patient need, and Pearl's insistence on distinguishing the two, and on making explicit what causal assumptions license using one as a stand-in for the other, is exactly the dissection this postmortem demands. His framework explains formally why Black patients generating lower costs at equal sickness is not noise but a systematic confound, unequal access and undertreatment intervening between true need and observed cost, that any naive proxy model will silently absorb as if it were signal. He did not personally investigate this healthcare case, keeping him just below the strongest carriers.

b. 1943 · stat-learning
90

Rubin's potential-outcomes framework for causal inference formalizes exactly the gap this scenario turns on: the algorithm needed to estimate a counterfactual quantity, how much healthcare a patient would require given equal access, but was trained on an observed quantity, cost actually incurred under unequal historical access, and his decades of work on the fundamental problem of causal inference explain precisely why substituting the observed for the counterfactual produces systematic bias whenever access itself is unequal across groups. His framework for re-deriving a model against a defensible causal target, rather than a convenient observational proxy, is exactly the fix this postmortem's final demand requires, keeping him among the very strongest carriers for this exact failure mode.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Differential Privacy

30 figures are scored on this problem. Draw it in a battle to see where you land.