AI History Battle

testing

Does the model fit at all?

It is 1900, and the whole apparatus of modern statistics is being built from scratch — including the very idea of asking whether data agree with a theory at all. You have five hundred observations and a fitted distribution you believe describes them. Devise a test of whether the data are genuinely consistent with that distribution, and be brutally clear about what 'reject' would and would not mean: rejection does not tell you the truth, only that this candidate is inadequate. Get it wrong and a wrong model is enshrined as fact, or an adequate one is discarded on a misread statistic. Calibration is everything — a goodness-of-fit test that is itself miscalibrated poisons every conclusion built on the model it blessed.

goodness-of-fitinfercalibration

Who this problem belongs to

The two figures whose methods fit it best, out of 49 in contention.

1857–1936 · early-stat
97

This problem is Pearson's 1900 paper described without his name. 'On the criterion that a given system of deviations from the probable...' introduced the chi-squared goodness-of-fit statistic: bin the five hundred observations, compare observed to expected counts under the fitted distribution, and refer the summed normalized squared deviations to a distribution he derived — arguably the founding act of testing whether data agree with a theory at all. He built it precisely because Victorian science enshrined curves by eye. The one deduction from a perfect score is calibration, the problem's own emphasis: Pearson got the degrees of freedom wrong when parameters are estimated, a miscalibration Fisher corrected in the 1920s over Pearson's fierce objection — proof that even the inventor could poison conclusions with a misread reference distribution. Still, no carrier's toolkit matches the problem more exactly.

1903–1987 · early-stat
92

This is nearly Kolmogorov's home turf. In 1933 — the same year as his measure-theoretic axiomatization of probability — he derived the exact limiting distribution of the supremum distance between the empirical distribution function and a hypothesized continuous distribution. The resulting Kolmogorov statistic is distribution-free under the null: one table calibrates the test for every continuous candidate, which is precisely the calibration guarantee the problem prizes. With five hundred observations his asymptotic distribution is already accurate, and the test requires nothing but sorting and a table. He would also be scrupulously clear that non-rejection is not confirmation. The only reasons he sits below Pearson here: the problem is set in 1900, three decades before his result, and his test assumes fully specified (not fitted-parameter) nulls — the estimated-parameter case needed later corrections.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hypothesis Testing

49 figures are scored on this problem. Draw it in a battle to see where you land.