testing
The lady and her teacups
It is 1935 in Cambridge, and a woman at a garden party claims she can taste whether the milk or the tea was poured into the cup first — a boast that becomes the founding parable of experimental testing. You present eight cups, four each way, in random order, and she must sort them. Work out, before she tastes a drop, how surprising a perfect or near-perfect score would be under the hypothesis that she is merely guessing, by enumerating the possible arrangements directly rather than leaning on any large-sample formula. Get it wrong and you credit a lucky guesser with a gift, or design a test too weak to reveal a real one — the whole logic of significance is built, cup by cup, on this exact enumeration.
Who this problem belongs to
The two figures whose methods fit it best, out of 45 in contention.
This problem is not an analogy for Fisher's work, it is Fisher's work: the Lady Tasting Tea episode at a 1935 Rothamsted garden party, which he formalized in The Design of Experiments (1935) as the founding example of a randomization test. Fisher insisted on enumerating the exact number of ways to choose four cups out of eight, C(8,4)=70, and computing the exact probability of a perfect or near-perfect sort under pure guessing, rather than invoking a normal approximation. That is precisely the enumeration-not-asymptotics discipline this problem demands. He also designed the randomized layout itself, tying experimental design to the test statistic's exact null distribution. No other candidate in this batch has a closer, more literal claim on this exact scenario than its own originator.
Writing as 'Student' at Guinness from 1908, Gosset pioneered exact small-sample inference precisely because brewery trial sizes were tiny and normal-theory approximations were unreliable; his 1908 paper derived the exact sampling distribution of the mean for small n, the ancestor move to Fisher's insistence on exact enumeration rather than asymptotic formulas. Gosset was Fisher's close correspondent and intellectual forerunner at the moment permutation-style small-sample reasoning displaced large-sample approximation in British statistics, and he explicitly argued for computing exact probabilities from small data rather than trusting the Gaussian limit. His toolkit transfers almost perfectly to enumerating the 70 arrangements of eight tea cups directly. He does not score above Fisher only because the tea-tasting logic and its explicit combinatorial framing are Fisher's own invention.
Fought here
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
45 figures are scored on this problem. Draw it in a battle to see where you land.