AI History Battle

testing

How big must the study be?

It is the planning meeting before any data exist, and a researcher must answer a question that decides whether the whole study is worth running: how many subjects are needed to have a good chance of detecting an effect of the size that would actually matter? Too few and the study is doomed to a non-significant result no matter what is true — a foregone waste; too many and you spend scarce funding and expose extra subjects for no added knowledge. Compute the required sample size from the target effect, the variability, and the error rates you will accept, and defend every assumption. Get it wrong and an underpowered study finds nothing and gets read, disastrously, as evidence that nothing is there.

powersample-sizedesign-before-data

Who this problem belongs to

The two figures whose methods fit it best, out of 50 in contention.

1894–1981 · early-stat
96

This problem is Neyman's own vocabulary made operational. His hypothesis-testing framework, built with Egon Pearson through the 1930s, is the first to treat Type II error and statistical power as first-class quantities to be controlled by design rather than discovered by accident, and the modern sample-size formula — solve for n given an effect size, a variance, a significance level, and a target power — is a direct descendant of the Neyman-Pearson lemma's logic about the most powerful test at a fixed error rate. Working at Berkeley from 1938, he insisted that the sample size be fixed and justified before data collection began, precisely the discipline this problem's planning-meeting framing demands. No other figure on this list built the specific mathematical apparatus this scenario asks the student to derive from as directly as Neyman did.

1890–1962 · early-stat
90

Fisher's program of randomized experimental design, developed at Rothamsted through the 1920s and codified in his 1935 Design of Experiments, made 'how big must the experiment be' a question to be answered mathematically before a single plot was sown, using expected effect sizes and error variance to justify replication counts. His analysis of variance gives the machinery for partitioning variability that any power calculation for a designed experiment ultimately rests on, and he was fiercely insistent that decisions about sample size and replication belonged in the planning stage, not as an afterthought. He scores just below Neyman because Fisher himself resisted the Neyman-Pearson framing of power and fixed error rates as the organizing principle, preferring his own significance-testing logic, even though his experimental-design practice anticipated exactly this problem's demands.

Fought here

Emmanuel Candes beat Rene Vidal 14–8

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hypothesis Testing

50 figures are scored on this problem. Draw it in a battle to see where you land.