AI History Battle

small-sample

Squeeze the estimator dry

It is 1945 in Cambridge, and a young statistician is asking a sharp question: given a small, expensive sample, is your obvious estimator wasting information you already paid for? There is a lower bound on the variance any unbiased estimator can achieve, and a procedure that, by conditioning on a sufficient statistic, provably improves a crude estimator toward it. Take a rough estimator built from a tiny sample and refine it to the minimum-variance one, and prove you cannot do better. Get it wrong and every scarce observation buys less certainty than it should — in small-sample work, where you can never simply gather more, leaving efficiency on the table is the difference between a conclusion and a shrug.

n smallinfersufficiency & efficiency

Who this problem belongs to

The two figures whose methods fit it best, out of 46 in contention.

1920–2023 · early-stat
98

This is Rao's own result. His 1945 paper for the Calcutta Mathematical Society, written while still a student under Fisher's shadow, derived the lower bound now called the Cramer-Rao bound and showed how conditioning a crude estimator on a sufficient statistic strictly reduces variance without introducing bias — the Rao-Blackwell construction he and Blackwell arrived at independently. He later built the fuller apparatus of information geometry to explain why the bound is tight only along certain curves in parameter space. Handed a rough small-sample estimator, refining it toward the minimum-variance unbiased estimator and proving optimality via the bound is exactly the sequence of moves he pioneered. The only reason it is not a perfect 100 is that in 1945 the geometric picture was still a decade off.

1919–2010 · stat-learning
95

Blackwell arrived at the same conditioning trick as Rao independently in 1947, and the theorem carries both names for good reason: given any unbiased estimator, conditioning on a sufficient statistic yields a new estimator that is never worse and strictly better unless the original was already a function of the sufficient statistic. His broader work in statistical decision theory, developed alongside Wald's minimax framework, gave the formal language of admissibility that makes 'you cannot do better' a provable claim rather than a slogan. Taking a crude small-sample estimator, conditioning it down to a sufficient statistic, and certifying it against the variance floor is precisely his procedure, executed by the man who proved it works and later applied it across games, dynamic programming, and Bayesian statistics.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hypothesis Testing

46 figures are scored on this problem. Draw it in a battle to see where you land.