It is the era when statistics is becoming an industrial process, and a colleague bursts in celebrating a p=0.03 — the one significant result out of twenty tests he ran. You must break the news gently and then repair the analysis: twenty independent looks at pure noise will hand you a 'significant' finding about two-thirds of the time, and his one survivor is very likely a mirage. Design the correction that controls the family-wide error rate without throwing away the power to detect effects that are genuinely there. Get it wrong and either the literature fills with false discoveries no one can replicate, or your correction is so brutal that real effects are dismissed — the multiplicity is the trap and the cure both.
Chose Distribution-free prediction intervals — wrong. Honest inference under weak assumptions was the one that fit.
Wasserman worked on large-scale multiple testing directly: with Christopher Genovese he developed false discovery control methods for functional neuroimaging, where every voxel is a hypothesis and the family runs to hundreds of thousands — the industrial extreme of the colleague's twenty. His stochastic-process view of FDR (treating the p-value threshold as an estimation problem with confidence guarantees) is exactly the kind of repair that controls family-wide error without Bonferroni's brutality. As the author of All of Statistics and a career-long bridge between statistics and machine learning, he can also explain the fix to the colleague at whatever level of rigor the situation demands. He arrives after Benjamini-Hochberg rather than originating the framework, so he sits below the inventors, but this is comfortably inside his research territory rather than adjacent to it.
Schölkopf systematized kernel methods in the 1990s — the kernel trick, support vector machines with Smola and colleagues, and later kernel mean embeddings — and one branch of that program does reach hypothesis testing: the kernel two-sample test (MMD, with Gretton and others) and kernel independence tests (HSIC) are genuine inferential tools his school built, so testing per se is not foreign territory. His causal-machine-learning turn also sharpened his sensitivity to how analyses mislead — confounding and selection are cousins of the multiplicity trap. But controlling error rates over families of tests was never his subject: the MMD line concerns constructing single powerful tests, not correcting twenty of them, and family-wise or FDR machinery appears in his work only as imported infrastructure. A methodologist near the border, contributing on the wrong side of it for this repair.
Battle #139 · 8/10/2026, 11:39:55 AM · this result is deterministic: the same two personas on this problem always resolve the same way.