high-dim
Inference after the search
It is 2015, and a quiet scandal runs through applied statistics: an analyst runs the lasso over ten thousand variables, keeps the dozen that survive, and then reports textbook p-values for those twelve — as if chosen before the data arrived. They were chosen by the data, and the classical guarantees are void; the selected effects are biased upward by the very act of winning the selection. Build inference that survives selection: either condition on the selection event and derive exact post-selection distributions, or construct knockoff variables — decoys statistically exchangeable with the originals — that let you control the fraction of false discoveries among the selected. Get it wrong and the genomics and neuroimaging literature keeps publishing winners'-curse effects that evaporate on replication, one confident dozen at a time.
Who this problem belongs to
The two figures whose methods fit it best, out of 62 in contention.
This is Barber's own founding result. With Candes, her 2015 paper 'Controlling the False Discovery Rate via Knockoffs' constructed exactly the decoy-variable procedure this problem describes: statistically exchangeable knockoff variables that let an analyst control the fraction of false discoveries among lasso-selected variables without needing to condition on the selection event analytically. That is a direct, exact answer to the problem's second bullet, developed in precisely this year on precisely this scandal. Her broader research program in conformal prediction and distribution-free inference after model selection extends and generalizes the same core insight: valid inference does not require pretending you didn't look at the data, if you build the right decoys or use the right conditioning. Dropped into 2015, she is not applying a toolkit to this problem; she built it.
Candes co-authored the knockoff-filter paper with Barber in 2015, making this problem essentially a restatement of his own published result, and his broader career-long expertise in high-dimensional statistics, incoherence conditions, and rigorous recovery guarantees under sparsity gives him the deepest possible technical grounding for both halves of this problem's demand: conditioning on selection events for exact post-selection distributions, and constructing exchangeable knockoffs to control false discovery. His compressed-sensing background also means he intuitively understands why the lasso's variable selection is itself a data-dependent random event that classical inference ignores at its peril. The only reason he is not scored above Barber is that the knockoff construction, the more novel and general half of their joint contribution, is more closely credited to her as lead author.
Fought here
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
62 figures are scored on this problem. Draw it in a battle to see where you land.