AI History Battle

It is the era when statistics is becoming an industrial process, and a colleague bursts in celebrating a p=0.03 — the one significant result out of twenty tests he ran. You must break the news gently and then repair the analysis: twenty independent looks at pure noise will hand you a 'significant' finding about two-thirds of the time, and his one survivor is very likely a mirage. Design the correction that controls the family-wide error rate without throwing away the power to detect effects that are genuinely there. Get it wrong and either the literature fills with false discoveries no one can replicate, or your correction is so brutal that real effects are dismissed — the multiplicity is the trap and the cure both.

multiplicityinfererror control
b. 1959
tapped · ask the professor
76

Chose Honest inference under weak assumptions — right call.

Wasserman worked on large-scale multiple testing directly: with Christopher Genovese he developed false discovery control methods for functional neuroimaging, where every voxel is a hypothesis and the family runs to hundreds of thousands — the industrial extreme of the colleague's twenty. His stochastic-process view of FDR (treating the p-value threshold as an estimation problem with confidence guarantees) is exactly the kind of repair that controls family-wide error without Bonferroni's brutality. As the author of All of Statistics and a career-long bridge between statistics and machine learning, he can also explain the fix to the colleague at whatever level of rigor the situation demands. He arrives after Benjamini-Hochberg rather than originating the framework, so he sits below the inventors, but this is comfortably inside his research territory rather than adjacent to it.

1928–1971
was tapped
15

Rosenblatt's perceptron (1958) was a trainable classifier, and his Principles of Neurodynamics (1962) reported extensive experiments — but his statistical practice was of the exploratory, demonstration-driven variety, and he became a cautionary tale about inference from enthusiastic evidence: press claims for the perceptron outran what his experiments established, and Minsky and Papert's 1969 analysis showed how much the celebrated results owed to unexamined limitations. He worked contemporaneously with Tukey's multiple-comparisons formalization but shows no engagement with it; error control in his world meant classification mistakes, not Type I rates over families of tests. Faced with the colleague's twenty tests, he had a Cornell psychologist's basic statistics training and genuine ingenuity, but neither the multiplicity machinery nor, historically, the demonstrated temperament for deflating an exciting borderline finding.

Head to head 10 over 1 battle
Read Wasserman Read Rosenblatt Leaderboard

Battle #172 · 8/10/2026, 11:42:33 AM · this result is deterministic: the same two personas on this problem always resolve the same way.