It is 2015, and a quiet scandal runs through applied statistics: an analyst runs the lasso over ten thousand variables, keeps the dozen that survive, and then reports textbook p-values for those twelve — as if chosen before the data arrived. They were chosen by the data, and the classical guarantees are void; the selected effects are biased upward by the very act of winning the selection. Build inference that survives selection: either condition on the selection event and derive exact post-selection distributions, or construct knockoff variables — decoys statistically exchangeable with the originals — that let you control the fraction of false discoveries among the selected. Get it wrong and the genomics and neuroimaging literature keeps publishing winners'-curse effects that evaporate on replication, one confident dozen at a time.
Chose Structure-constrained high-dimensional inference — right call.
Wasserman's work bridging classical statistics and modern machine learning, and his insistence on honest, well-calibrated inference in complex high-dimensional settings, gives him direct, rigorous facility with this problem's core demand: that a lasso-selected dozen variables cannot be treated as if chosen before the data arrived. His statistics-ML bridge gives him fluency with both the classical hypothesis-testing literature this problem's scandal violates and the modern selective-inference and knockoff methods that fix it. He is not among the strongest carriers because the specific founding constructions, exact post-selection conditioning and the knockoff filter, were developed by Barber, Candes, and Taylor's collaborations rather than his own research program, though he has written clearly and rigorously about exactly this problem.
Breiman's foundational work on classification and regression trees, bagging, and random forests, alongside his 'two cultures' critique of statistical modeling, gives him general, rigorous statistical facility with the dangers of overfitting and false confidence in models fit to the same data used to select them, a real but general connection to this problem's core scandal. He has no direct engagement with the specific post-selection conditional-distribution and knockoff-filter constructions that Barber, Candes, and Taylor developed as this problem's precise mathematical answer; his own most prominent contributions address tree-based ensemble methods rather than lasso-specific selective inference. That mix of genuine conceptual adjacency and real distance from the specific selective-inference literature is what keeps Leo Breiman solidly in this batch's middle tier.
Battle #90 · 8/10/2026, 11:37:02 AM · this result is deterministic: the same two personas on this problem always resolve the same way.