AI History Battle
p = 20,000, n = 200 regression

It is the genomic era, and a microarray hands you twenty thousand gene-expression measurements on just two hundred patients — vastly more predictors than people, a regime where classical regression simply dissolves, fitting the noise perfectly and generalizing not at all. Somewhere in those twenty thousand genes, a handful actually drive the disease. Find that handful, and — harder — say honestly how confident you are that each is signal and not one of the countless noise genes that will, by sheer number, mimic a real effect. Get it wrong and biologists chase phantom genes for years, or a real drug target is buried under false positives. In the p ≫ n world, sparsity and honest selection are the only things standing between you and self-deception.

high-dim sparseinfer+predictregularization
b. 1974
tapped · ask the professor
51

Chose Causal-state reconstruction — wrong. The misspecification audit was the one that fit.

Shalizi is a statistician steeped in model selection, complexity, and the philosophy of honest inference — exactly the mindset the problem demands when twenty thousand genes invite over-fitting and self-deception. He writes incisively on cross-validation, regularization, the dangers of selection without correction, and when predictive models can be trusted, and he knows the high-dimensional literature well. He scores around midpoint because his command of the conceptual tools — bias-variance, penalization, multiple-testing pitfalls — is strong and directly on-point, but his own research contributions lean toward computational mechanics and complexity rather than inventing sparse-recovery or FDR-control methods for genomics. He would correctly diagnose the traps and pick sound tools, yet the defining instruments of this problem are others'. A well-equipped generalist, mid-batch.

b. 1987
was tapped
18

Goodfellow invented generative adversarial networks — a breakthrough for generating realistic high-dimensional samples like images, given large training sets. That is orthogonal to the problem: selecting a sparse handful of driver genes from twenty thousand on two hundred patients is a supervised, inference-heavy, small-sample task with no generative-modeling need and far too little data to train adversarial networks. His broader deep-learning expertise assumes the big-n regime this problem denies. He scores near the floor: a landmark researcher whose signature tool solves a different kind of problem entirely. GANs offer no mechanism for honest variable selection or false-discovery control, and the sample size forecloses deep generative modeling, leaving his instruments essentially inapplicable to the specific genomic selection-and-inference challenge posed here.

Head to head 01 over 1 battle
Read Shalizi Read Goodfellow Leaderboard

Battle #94 · 8/10/2026, 11:37:24 AM · this result is deterministic: the same two personas on this problem always resolve the same way.