AI History Battle

testing

When the bell curve won't hold

It is 1947, and two small groups of measurements must be compared, but the data are skewed and lumpy — nothing like the normal curve the t-test quietly assumes — so its p-value cannot be trusted here. You need a test that compares the two groups using only the ranks of the values, not their raw magnitudes, so its validity does not depend on any distributional shape at all. Derive the distribution of the rank statistic under the null and defend why a distribution-free test buys robustness at only a small cost in power when normality actually holds. Get it wrong and you either trust a t-test whose assumptions the data violate and declare a false difference, or discard a real effect for lack of a valid test.

nonparametricrank testdistribution-free

Who this problem belongs to

The two figures whose methods fit it best, out of 42 in contention.

1915–2000 · early-stat
93

Tukey's lifelong commitment to robust statistics and distribution-free methods, developed from the 1940s onward and crystallized in his later work on exploratory data analysis, is built precisely around this problem's premise: real data are skewed and lumpy, and a competent statistician needs tools that do not silently assume away that reality. His extensive contributions to rank-based and robust estimation, including his work on trimmed means and influence functions, address exactly the tradeoff the problem demands defending, that distribution-free tests buy robustness at only a small cost in power when normality actually holds. His practical, skeptical instinct for testing assumptions before trusting a parametric procedure is precisely the disposition this 1947 problem requires. He is marked just short of perfect only because the specific rank-sum derivation this problem describes predates his own major publications by several years.

b. 1940 · stat-learning
87

Bickel's foundational work on robust and adaptive estimation, developed at Berkeley from the 1970s onward, directly engages this problem's core question: how much statistical efficiency must be sacrificed to gain validity when the data's true distribution is unknown or non-normal, and how an adaptive procedure can approach full parametric efficiency without assuming a specific distributional form. His semiparametric efficiency theory gives rigorous mathematical language for exactly the small-power-cost tradeoff this problem asks the analyst to defend. His deep grounding in the mathematical foundations connecting classical rank-based nonparametric tests to modern semiparametric theory makes him an unusually strong match. He is marked slightly below Tukey because his major contributions arrive decades after this 1947 scene, refining rather than originating the rank-test tradition the problem is set inside.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hypothesis Testing

42 figures are scored on this problem. Draw it in a battle to see where you land.