testing
Two proportions, tiny cells
It is the 1930s, and two clinics report cure rates you must compare: three of eleven at one, eight of twelve at the other. The chi-squared approximation everyone reaches for was built for large samples, and with cells this small its p-values are folklore — comforting and wrong. You need exact inference: enumerate the possible tables directly and compute the true probability of a split this extreme under chance. Defend why the approximation fails and the exact calculation must replace it. Get it wrong and a clinic is credited with a cure it never delivered, or a real difference between treatments is waved away as noise — at this cell size, only the exact combinatorics tell the truth, and the shortcut lies.
Who this problem belongs to
The two figures whose methods fit it best, out of 51 in contention.
This problem is Fisher's own invention answered in his own decade. By the mid-1930s Fisher had published the exact test for 2x2 contingency tables — conditioning on the fixed margins and computing the hypergeometric probability of every table as extreme as the one observed — precisely because Pearson's chi-squared approximation collapses when expected cell counts are small. Statistical Methods for Research Workers carried the method to working scientists, and The Design of Experiments (1935) framed exact permutation reasoning with the lady-tasting-tea example. Given three of eleven versus eight of twelve, Fisher would fix the margins, enumerate the handful of admissible tables, sum the tail probabilities, and defend the conditioning argument against all comers. No anachronism, no gap: the 1930s toolkit and the problem match exactly.
Gosset spent his career at Guinness doing exact inference on samples the size of a brewing batch — often four to ten observations — and his 1908 paper 'The Probable Error of a Mean' exists because he refused to trust large-sample normal theory at those sizes. His signature t-distribution addresses means rather than proportions, so the specific hypergeometric enumeration here is Fisher's, not his; but the diagnosis — that asymptotic p-values are folklore when cells are tiny — is the founding insight of Gosset's entire program, articulated decades before most statisticians accepted it. Working by hand and by simulation (he famously shuffled cards to check his sampling distributions), he would verify the exact calculation empirically. He corresponded with both Pearson and Fisher and understood exactly where the approximation breaks.
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
51 figures are scored on this problem. Draw it in a battle to see where you land.