AI History Battle

testing

Five sigma at the collider

It is 2012 at CERN, and a bump is rising in the data where the Higgs boson should be — but particle physics will not call it a discovery until the signal reaches five sigma, a false-positive standard thousands of times stricter than the rest of science uses. Justify that stringency, then confront the subtlety that makes it necessary: the "look-elsewhere effect," because the bump could have appeared at many masses, and a threshold impressive at one fixed spot is far less so when you scanned a whole range. Correct for it honestly. Get it wrong and you announce the discovery of the century on a fluctuation, or — erring the other way — sit on a genuine particle for fear of the multiplicity you failed to quantify.

extreme significancemultiplicitylook-elsewhere

Who this problem belongs to

The two figures whose methods fit it best, out of 41 in contention.

1890–1962 · early-stat
88

Fisher gave the entire physics community the vocabulary it uses at CERN: his 1925 development of significance testing and the p-value as a measure of surprise against a null hypothesis is the direct ancestor of the five-sigma standard, which is simply an extraordinarily stringent p-value threshold expressed in standard-deviation units. His insistence that significance be judged relative to a precisely specified null and that the analysis be planned with discipline anticipates the rigor a Higgs-search analysis demands. But Fisher never confronted anything resembling the look-elsewhere effect at this scale — searching a continuous range of possible masses for a bump — since his agricultural and biological applications rarely involved scanning thousands of effectively independent tests, keeping him just below the two figures who built the actual error-control machinery this problem requires.

1894–1981 · early-stat
86

Neyman's hypothesis-testing framework, built with Egon Pearson through the 1930s, gives five-sigma its precise operational meaning — a fixed, pre-specified significance level chosen so stringently that the long-run rate of false discovery across all such claims is vanishingly small, exactly the logic CERN's physicists invoke to justify their threshold. His framework of composite hypotheses and his insistence on controlling error rates across repeated applications of a test is the direct theoretical ancestor of correcting for the look-elsewhere effect, which is fundamentally a multiple-comparisons problem across a scanned mass range. He does not edge out Fisher and Efron only because the specific trials-factor correction physicists use to handle a continuously scanned search region was formalized by later statisticians building explicitly on his foundational error-rate framework.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hypothesis Testing

41 figures are scored on this problem. Draw it in a battle to see where you land.