small-sample
Ten patients, one rare disease
It is 1908, and a physician has tried a new therapy on ten patients suffering a rare condition; six improved, where the historical rate is only 30%. Ten is all there will ever be this year — the disease is rare, the patients scattered, and each case took months to find. Decide what can honestly be concluded from six-in-ten against a thirty-percent baseline, and state every caveat that a small numerator forces on you. Use exact methods, not large-sample approximations that lie at this scale. Get it wrong and a useless treatment enters practice, or a real cure is abandoned — either way patients pay for a statistician's overconfidence. n=10 is the entire evidence base.
Who this problem belongs to
The two figures whose methods fit it best, out of 41 in contention.
This is Gosset's problem in the very year he owned it. In 1908, publishing as 'Student' because Guinness forbade employees from publishing, he gave the world 'The Probable Error of a Mean' -- exact small-sample inference invented precisely because brewery experiments yielded handfuls of observations and large-sample formulas lied at that scale. His t-distribution concerns means of normal samples rather than binomial counts, but the exact binomial tail for six-in-ten against p=0.3 is hand-computable arithmetic he performed routinely, and his temperament fits the brief: an applied scientist who treated overconfidence as an operational cost measured in wasted barley and bad policy. He would compute the exact probability, refuse the normal approximation outright, and phrase the conclusion with a brewer's practical humility about what ten cases permit.
Laplace solved essentially this calculation a century before 1908. His rule of succession and his analysis of the Paris birth-ratio data show him computing exact posterior probabilities for a binomial proportion by hand, integrating what we now call the beta posterior with the analytical machinery he built for celestial mechanics. Given six successes in ten trials, he would report the posterior probability that the true rate exceeds the 30% baseline -- a direct, exact, finite-sample answer requiring no approximation, though he derived the normal approximation too and knew exactly when it failed. The era gaps run both ways: he predates any notion of confounded historical controls or selection of patients, and his default uniform prior ignores the baseline information; but no one in history was better at exact hand computation of precisely this quantity.
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
41 figures are scored on this problem. Draw it in a battle to see where you land.