AI History Battle

testing

Does the extra parameter earn its keep?

It is 1938, and a recurring question in model-building needs a general answer: you have a simple model and a richer one that contains it, and the richer one always fits the data at least as well — so when is its improvement real rather than the inevitable reward of extra freedom? Derive a test based on the ratio of maximized likelihoods, show that twice its logarithm follows a known distribution under the null, and connect it to the twin tests built from the curvature of the likelihood and from its slope at the simpler model. Get it wrong and you either keep adding parameters that only fit noise, or reject a genuine refinement — the whole discipline of nested model comparison rests on knowing what improvement chance alone supplies.

likelihood-rationested modelsasymptotics

Who this problem belongs to

The two figures whose methods fit it best, out of 48 in contention.

1890–1962 · early-stat
96

This problem is Fisher's own research program stated as a question. His maximum likelihood theory, developed through the 1910s and 1920s and consolidated in his 1922 and 1925 papers, made the likelihood function the single organizing object for comparing how well nested models explain data, and his notion of sufficiency and efficiency underlies why a richer model's gain in likelihood has a predictable null distribution when the simpler model is actually true. He did not personally derive Wilks's 1938 result that twice the log-likelihood-ratio is asymptotically chi-squared, but every piece of machinery that result depends on — the likelihood function, its regularity conditions, its large-sample behavior — is his construction. The score just short of a perfect mark reflects that the precise asymptotic theorem belongs to a student of his tradition rather than to Fisher himself.

1894–1981 · early-stat
91

Neyman, with Egon Pearson, built the general logic this problem calls for: the 1933 Neyman-Pearson lemma proves that among all tests at a fixed significance level, the likelihood-ratio test is the most powerful, which is exactly the justification needed for why comparing maximized likelihoods is the principled way to ask whether a richer model's fit is real. His framework of Type I and Type II errors gives the vocabulary for 'when is improvement real versus the reward of extra freedom' its precise operational meaning. Working at Berkeley from 1938 onward, he extended this apparatus into composite hypothesis testing, the exact regime nested-model comparison lives in. He scores just below Fisher because the asymptotic chi-squared distribution of the statistic itself was Wilks's specific contribution, building on but distinct from the Neyman-Pearson lemma.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hypothesis Testing

48 figures are scored on this problem. Draw it in a battle to see where you land.