AI History Battle

testing

The p-value reckoning

It is 2016, and a professional statistical society does something unprecedented: it issues a public warning that the p-value, the field's most-used number, has been routinely misunderstood and abused into a replication crisis. You are handed a "significant" published finding and asked to interrogate it: how many analyses were silently tried before this one crossed the threshold, whether the hypothesis was fixed before the data or chosen after, and what the p-value does and does not say about the truth of the claim. Rebuild the analysis so its uncertainty is honest. Get it wrong and you add another irreproducible result to a literature already drowning in them — the crisis is not that the arithmetic is hard, but that a number this seductive invites self-deception.

replicationp-hackingmeta

Who this problem belongs to

The two figures whose methods fit it best, out of 56 in contention.

b. 1965 · stat-learning
97

Gelman is perhaps the single most prominent public voice on exactly this crisis. From his Columbia blog and decades of writing on statistical practice, he coined the 'garden of forking paths' to describe how researcher degrees of freedom — the many silent choices made before a result crosses a threshold — inflate false-positive rates without any deliberate p-hacking at all. His work with Eric Loken formalized this critique well before the 2016 ASA statement, and his broader push for preregistration, effect-size skepticism, and Bayesian hierarchical modeling as replacements for naive significance testing directly answers this problem's demand to interrogate a published finding's hidden analytic flexibility. Few statisticians have spent more of their career diagnosing precisely this failure mode in precisely this way, making him about as close to a perfect match as the roster offers.

b. 1959 · stat-learning
92

Wasserman's dual identity as a rigorous mathematical statistician and an unusually candid public commentator on statistical practice, developed through his Normal Deviate blog and his textbook All of Statistics, makes him a natural fit for interrogating a suspect published finding. He has written pointedly about the abuse and misunderstanding of p-values, the tension between statistical and practical significance, and the need for statisticians to take replication failures seriously rather than defend the field reflexively. His comfort bridging classical frequentist theory and modern statistical-learning perspectives lets him diagnose both the technical misuse (multiple comparisons, p-hacking) and the conceptual confusion (mistaking a p-value for the probability the null is true) this problem describes. He is marked just below Gelman because his commentary, while sharp, was less central to the specific 2016 ASA reckoning than Gelman's sustained public campaign.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hypothesis Testing

56 figures are scored on this problem. Draw it in a battle to see where you land.