high-dim
Does your pipeline reproduce?
It is 2015 at Berkeley, and the replication crisis has reached data science: published findings from high-dimensional pipelines — genomics hits, neuroimaging correlates, precision-medicine signatures — keep dissolving when anyone reruns them, because every pipeline hides dozens of reasonable-seeming analytic choices, and the published result is one path through that garden of forking paths. Take a published high-dimensional analysis and stress it: perturb the data by resampling, perturb the analytic choices — thresholds, normalizations, model families — and report which findings survive the ensemble. Then go further: formalize stability as a criterion for trustworthy inference, on par with predictive accuracy, not an afterthought. Clinical decisions and drug-target investments are being built on findings of exactly this kind. A result that cannot survive its own reasonable alternatives was never a result.
Who this problem belongs to
The two figures whose methods fit it best, out of 59 in contention.
This is Yu's own program, stated almost verbatim: her 'veridical data science' and the PCS (predictability-computability-stability) framework, developed at Berkeley through the 2010s, exist specifically to formalize stability as a criterion for trustworthy inference alongside predictive accuracy. She built perturbation-based diagnostics — resample the data, perturb the analytic pipeline's choices, and report which findings survive — as a working methodology, not a metaphor, applied to genomics and neuroscience pipelines of exactly this kind. Her era gap is essentially zero; she was writing this argument in 2015 while Berkeley colleagues were living the replication crisis in real time. No carrier's toolkit maps more directly onto 'formalize stability, stress the pipeline, report what survives' than the person who invented the framework for it.
Shalizi's writing across the 2000s-2010s — statistics blog, textbook drafts, papers on model misspecification — is a sustained methodological critique of exactly this failure mode: pipelines that look rigorous but encode arbitrary analytic choices dressed as objectivity. His computational mechanics work insists on asking whether a model's structure is identifiable from data at all, which is the deeper question behind 'does this survive perturbation.' He has the statistical machinery (nonparametrics, model selection critique) and the polemical clarity to formalize forking-paths problems. What he lacks relative to Yu is a single canonical algorithmic framework bearing his name; his contribution is diagnostic and critical rather than a deployed stability pipeline, so he ranks just below her.
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
59 figures are scored on this problem. Draw it in a battle to see where you land.