testing
Which pairs really differ?
It is 1953, and an analysis of variance has just told you that five fertilizer treatments are not all equal — but that verdict names no winner, and the moment you start comparing the treatments two at a time, you are running ten tests and inflating your false-positive rate all over again. Design a procedure for these post-hoc comparisons that controls the error rate across the entire family of pairwise contrasts, so that the specific differences you declare significant are trustworthy. Get it wrong and you either declare pairs different on the strength of the largest chance gap among many, crowning a fertilizer that only got lucky, or you correct so harshly that the real difference the ANOVA promised exists can never be located.
Who this problem belongs to
The two figures whose methods fit it best, out of 40 in contention.
This problem is dated to 1953 for a precise reason: that is the year Tukey circulated his honestly significant difference procedure, the first rigorous solution to exactly this dilemma of an ANOVA declaring an overall effect while naming no specific winner. His HSD test uses the studentized range distribution to set a single critical value for all pairwise comparisons simultaneously, controlling the family-wise error rate across the entire set of contrasts rather than letting it inflate test by test, which is precisely the mathematical object this scenario asks the student to design. He developed it working through problems in agricultural and industrial experimentation very close to this fertilizer-trial scenario. No other figure on this list built the specific procedure this problem is named for, in the exact year the problem specifies.
Fisher's analysis of variance, developed at Rothamsted through the 1920s, is the very tool that hands down the verdict this problem opens with — that the five treatments are not all equal — and his least significant difference procedure, developed as a follow-up test after a significant F-statistic, was the first widely used attempt at exactly the post-hoc pairwise comparison this problem demands. But Fisher's LSD test, unlike Tukey's later HSD, does not fully control the family-wise error rate across all pairs simultaneously, a gap that motivated Tukey's correction in the first place; Fisher was also famously resistant to the very idea that correcting for multiple comparisons was necessary, arguing the ANOVA's significance already licensed follow-up tests, a position later statisticians substantially revised.
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
40 figures are scored on this problem. Draw it in a battle to see where you land.