AI History Battle

causality

The therapy the trial reversed

It is July 2002, and the Women's Health Initiative has just stopped a trial early: hormone replacement therapy, prescribed to millions on the strength of consistent observational studies showing cardiac benefit, is raising heart risk in the randomized data, not lowering it. Two literatures, same drug, opposite signs. Perform the autopsy: identify how healthy-user confounding — the women who chose HRT were richer, fitter, better doctored — manufactured a protective association out of selection, and why statistical adjustment for recorded covariates could not remove what was never recorded. Then state what observational evidence can and cannot support going forward. Get it wrong and the next plausible therapy rides the same confounded evidence into millions of prescriptions.

confounding by indicationobservational vs randomizedepidemiology at scale

Who this problem belongs to

The two figures whose methods fit it best, out of 57 in contention.

b. 1943 · stat-learning
93

Rubin's potential-outcomes framework, developed from the early 1970s onward, gives the precise vocabulary for exactly what went wrong in the hormone-replacement-therapy literature: healthy-user confounding is a case where the treatment-assignment mechanism, women who chose HRT were systematically richer, fitter, and better doctored, correlates with the outcome through paths no recorded covariate captures, so adjustment cannot recover the missing randomization. His decades of work distinguishing what observational data can and cannot support, and his insistence that a well-specified causal question requires an explicit assignment mechanism, is the direct intellectual tool for performing the autopsy the problem demands. This is not adjacent theory transferred sideways; it is the exact causal-inference apparatus built to diagnose precisely this failure mode, which is why the score sits near the ceiling.

1924–2022 · early-stat
82

Cox's proportional hazards model, introduced in 1972, became and remains the standard statistical tool for exactly this class of problem, estimating the effect of a treatment like hormone therapy on a time-to-event outcome like cardiac disease from observational cohort data, and his broader career-long rigor about the design and limits of epidemiological studies gives him deep authority on precisely when such models can and cannot be trusted. His insistence on careful specification of what a hazard ratio actually estimates, and under what assumptions about confounding, speaks directly to the problem's demand to state what observational evidence can and cannot support going forward. The gap to the very top is that Cox's own applications were general epidemiological methodology rather than this specific 2002 Women's Health Initiative reversal.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Causal Inference Bayesian Networks

57 figures are scored on this problem. Draw it in a battle to see where you land.