It is the genomic era, and a microarray hands you twenty thousand gene-expression measurements on just two hundred patients — vastly more predictors than people, a regime where classical regression simply dissolves, fitting the noise perfectly and generalizing not at all. Somewhere in those twenty thousand genes, a handful actually drive the disease. Find that handful, and — harder — say honestly how confident you are that each is signal and not one of the countless noise genes that will, by sheer number, mimic a real effect. Get it wrong and biologists chase phantom genes for years, or a real drug target is buried under false positives. In the p ≫ n world, sparsity and honest selection are the only things standing between you and self-deception.
Chose Sparse recovery — right call.
Nowak works on sparse recovery, compressed sensing, and active learning — the signal-processing side of exactly this regime. His research on sparse signal reconstruction and on sample-complexity limits for recovering a few nonzero coefficients from underdetermined systems maps directly onto finding a handful of driver genes among twenty thousand from two hundred samples. He studied when correlated or structured sparsity helps, relevant to co-regulated genes, and his active-learning work addresses spending a limited measurement budget wisely, echoing the n = 200 constraint. He is strong on recovery guarantees and thresholds; slightly behind the very top tier because his emphasis is signal-processing recovery rather than the biostatistical false-discovery-control apparatus for honest gene calling. Still, his toolkit is a genuine, era-appropriate fit for the sparse-regression core.
Scholkopf systematized kernel methods and regularization theory — support vector machines with L2 penalties handle p >> n naturally by working in feature space and controlling capacity via margins, and SVMs were in fact widely applied to microarray classification. His representer-theorem and regularization framing directly address overfitting when features vastly outnumber samples. His later causal-inference work bears on distinguishing driver genes from correlated passengers, a real concern in the problem. He scores above midpoint because kernel/regularization machinery genuinely applies, but below the sparsity specialists: standard kernel methods give dense, not sparse, solutions, so they classify well but do not natively return a short interpretable gene list with per-gene significance. The 'find the handful and say how confident' demand is not the SVM's home turf, tempering the score.
Battle #150 · 8/10/2026, 11:40:29 AM · this result is deterministic: the same two personas on this problem always resolve the same way.