AI History Battle

testing

The bomber that came home

It is 1943, and you sit with the Statistical Research Group in Manhattan, staring at maps of returning bombers peppered with bullet holes — dense on the wings and fuselage, sparse on the engines. The generals want armor added where the holes cluster. That instinct is exactly backwards, and lives ride on your seeing why: you are looking only at the planes that made it home. The engine-hit bombers are not in your data because they are at the bottom of the Channel. Decide where the armor truly belongs and formalize the selection effect that makes the naive answer lethal. Get it wrong and you armor the wrong plates, and more crews don't come back — the missing data is the whole message.

selection biasinfersurvivorship

Who this problem belongs to

The two figures whose methods fit it best, out of 52 in contention.

1902–1950 · early-stat
98

This is Wald's own problem: he was at the Statistical Research Group at Columbia in 1943, and his memoranda on estimating aircraft vulnerability from damage on surviving planes are the canonical treatment. His move was to model the survivorship process explicitly — condition on the plane returning, treat hit probabilities per section as parameters, and back out the vulnerability of each section from the holes you do not see on survivors. He did this with the era's tools: conditional probability, likelihood-style reasoning, and careful bookkeeping over hypothetical downed aircraft, no computers required. His broader decision-theoretic framing (loss functions, minimax) matches the armor-allocation decision exactly. No later toolkit — causal graphs, missing-data theory — improves materially on what he actually wrote for this exact dataset. The problem is named for him in every methods course since.

b. 1936 · stat-learning
90

Pearl's toolkit, built from Bayesian networks in the 1980s through the do-calculus of the 1990s, is the modern formalization of exactly this trap. Selection bias in his framework is conditioning on a collider or on a selection node: the observed sample is generated by an arrow from engine-hit into survival, and analyzing survivors conditions on that node. His later work with Bareinboim on recoverability from selection-biased data gives explicit graphical criteria for when the generals' quantity — hit vulnerability in the full fleet — is or is not identifiable from returned planes plus assumptions. The anachronism cuts one way only: his machinery arrives forty years after 1943, but it is precisely the general theory of which Wald's memo is the founding special case. He would draw the graph and the fallacy would be visible in one diagram.

Fought here

Andrew Gelman beat Rudolf Kalman 68–31

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hypothesis Testing

52 figures are scored on this problem. Draw it in a battle to see where you land.