testing
One test or twenty?
It is the era when statistics is becoming an industrial process, and a colleague bursts in celebrating a p=0.03 — the one significant result out of twenty tests he ran. You must break the news gently and then repair the analysis: twenty independent looks at pure noise will hand you a 'significant' finding about two-thirds of the time, and his one survivor is very likely a mirage. Design the correction that controls the family-wide error rate without throwing away the power to detect effects that are genuinely there. Get it wrong and either the literature fills with false discoveries no one can replicate, or your correction is so brutal that real effects are dismissed — the multiplicity is the trap and the cure both.
Who this problem belongs to
The two figures whose methods fit it best, out of 45 in contention.
This is Tukey's home problem by name: his 1953 manuscript 'The Problem of Multiple Comparisons' — circulated for decades before formal publication — defined the field, and his honestly significant difference procedure gave simultaneous pairwise intervals that control the family-wise error rate while remaining usable at the bench. Just as important, his exploratory-versus-confirmatory distinction diagnoses the colleague's sin precisely: twenty looks at the data are exploration, and exploration may suggest but never confirm. Tukey would neither accept the p=0.03 nor bludgeon it with a naive Bonferroni; he built procedures calibrated to the actual family of comparisons, preserving power for real effects. Working mid-century without FDR machinery, he nonetheless framed error rates per family, per experiment, and per comparison — the exact vocabulary this repair requires. Few carriers own a problem this completely.
Efron spent his late career on exactly this, at industrial scale: the microarray era handed statisticians tens of thousands of simultaneous tests, and his empirical-Bayes response — local false discovery rates, estimating the null distribution from the data itself — became the standard treatment, codified in his 2010 monograph Large-Scale Inference. Where twenty tests strain Bonferroni, Efron's insight is that a family of tests is itself data: the ensemble reveals how many nulls there are and how severe the selection effect is. His bootstrap (1979) also supplies resampling-based stepdown procedures that respect dependence among the tests, recovering power a worst-case correction throws away. The one era gap runs in his favor — he arrived after Tukey posed the problem and helped build its modern, power-preserving cure. Direct, deep applicability.
Fought here
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
45 figures are scored on this problem. Draw it in a battle to see where you land.