Andrew Gelman
Hierarchical Bayes; Stan; applied Bayesian workflow
Played by Gina
Strongest on
Battles
When you can't randomize L Richard Bellman
Catch the process the moment it drifts W John Santerre
Peeking at the trial L Frederick Jelinek
The imitation game, scored W Frederick Jelinek
A hundred sensors for a city's water L Rudolf Kalman
Best answer before the buzzer W Rudolf Kalman
The bomber that came home L Rudolf Kalman
The variables that must be whole W Rudolf Kalman
The adaptive dose-finder L Rudolf Kalman
Predict the ore grade underground L Rudolf Kalman
What happened first?
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
Life and career
Andrew Gelman is probably the most-read statistician alive, and not because of his papers. For close to two decades he has posted, most days, to a blog with the unglamorous title *Statistical Modeling, Causal Inference, and Social Science*, in which he works through applied problems, argues with commenters, and — the part that made him famous outside statistics — takes apart published research that does not hold up. He has been, by some distance, the most persistent public voice in the replication crisis, and he has been that while also writing the standard graduate textbook on Bayesian analysis and co-founding the probabilistic programming system that a large fraction of applied Bayesians now use.
He studied mathematics and physics at MIT and took his doctorate in statistics at Harvard under Donald Rubin — a lineage that matters enormously for reading his work. Rubin's program was the potential-outcomes framework for causal inference, multiple imputation for missing data, and a Bayesian sensibility that treated modeling as an applied craft rather than a philosophical commitment. Gelman inherited all three and pushed each further.
After a few years on the faculty at Berkeley he moved to Columbia, where he has been ever since, with appointments in statistics and political science and a long-running applied statistics center. The political science half is not ornamental. A great deal of his most cited applied work is about American elections and public opinion: the relationship between income, state, and voting behavior, which he laid out in *Red State, Blue State, Rich State, Poor State*, and the estimation of state- and subgroup-level opinion from national surveys, which produced one of his most widely used methods. He has also worked on toxicology, pharmacology, education policy, and — memorably — a widely cited analysis of racial disparities in New York City police stop-and-frisk practices.
He received the COPSS Presidents' Award, the profession's top honor for a statistician under 40. But the honor that best captures his position in the field is informal: an enormous number of working scientists in the social and biomedical sciences have changed how they analyze data because Gelman argued with them, in public, at length, and would not let it go.
Key contributions
**Hierarchical models and partial pooling.** The core technical idea Gelman has spent his career elaborating is this: when data come in groups — patients within hospitals, students within schools, counties within states — you have two bad options and one good one. Pool everything (ignore the groups, and you throw away real variation) or estimate each group separately (no pooling, and small groups get wildly noisy estimates). The good option is to model the group-level parameters as themselves drawn from a distribution whose parameters you also estimate. The result is *partial pooling*: each group's estimate is shrunk toward the overall mean by an amount determined by its own sample size and by the estimated between-group variance. Hospitals with 30 cases get shrunk hard; hospitals with 3,000 barely move.
*Bayesian Data Analysis*, with Carlin, Stern, Rubin and later Dunson and Vehtari, codified this as the default treatment of grouped data, and its worked examples — the eight schools, radon measurements by county — became the canonical teaching cases. *Data Analysis Using Regression and Multilevel/Hierarchical Models*, with Jennifer Hill, took the same material to applied researchers who needed to run it rather than derive it.
Two refinements are worth knowing. First, his 2006 work on **prior distributions for variance parameters** showed that the then-standard "noninformative" inverse-gamma priors are badly behaved when the between-group variance is near zero — precisely the regime where the prior matters most — and proposed half-Cauchy and folded-noncentral-t alternatives, which are now standard. Second, **weakly informative priors** for logistic regression, which prevent the separation pathology that sends maximum-likelihood coefficients to infinity, while encoding only the mild claim that effects on the logit scale are not astronomically large.
**MRP.** Multilevel regression and post-stratification estimates subgroup quantities — opinion in a small state, say — by fitting a multilevel model with demographic and geographic predictors and then reweighting the model's predictions to match known population counts. It lets you get respectable state-level estimates from a national sample that contains twenty people from Wyoming, and it famously allowed election forecasts from wildly non-representative online panels. It is now the standard method in survey-based political science.
**Diagnostics and workflow.** The **Gelman-Rubin R̂ statistic**, from his work with Rubin, runs multiple MCMC chains from dispersed starting points and compares between-chain to within-chain variance; values near one indicate the chains have mixed. It is the single most-used convergence diagnostic in Bayesian computation. **Posterior predictive checks**, developed with Meng and Stern, formalize a habit good analysts already had: simulate replicate datasets from the fitted model and ask whether they look like the data you actually saw, in whatever respect you care about. If your model cannot reproduce a salient feature of the data, it does not matter what its posterior says.
More broadly, Gelman has argued for **Bayesian workflow** as a distinct object of study: model building is iterative, models are fitted, checked, expanded, and refitted, and this loop needs to be described honestly rather than hidden.
**Stan.** With Bob Carpenter, Matt Hoffman, and a growing team, Gelman co-founded Stan, a probabilistic programming language with automatic differentiation and Hamiltonian Monte Carlo at its core. Hoffman and Gelman's **No-U-Turn Sampler** removed HMC's most painful tuning parameter — the trajectory length — by adaptively extending the trajectory until it starts doubling back. Stan made high-dimensional, correlated, hierarchical posteriors routinely samplable by people who are not MCMC specialists, and its influence on applied Bayesian practice is hard to overstate.
**Methodological critique.** Gelman and Loken's "garden of forking paths" made a sharp conceptual contribution to the replication debate: you do not need deliberate p-hacking to get a meaningless significant result. It suffices that the analysis you would have run depends on the data you happened to see — the researcher makes one defensible choice, honestly, but many other defensible choices were available and the significance is conditional on the path taken. Relatedly, his work with Carlin on **Type S and Type M errors** reframed underpowered studies: the danger is not just failing to detect an effect, it is that any effect you *do* detect will be exaggerated in magnitude and may well have the wrong sign. He has also argued, with Hill and Yajima, that multilevel modeling handles multiple comparisons more sensibly than post-hoc corrections, because shrinkage already accounts for the multiplicity.
In battle
Gelman has the strongest and broadest profile in this cohort by a comfortable margin: mean 51.3, median 52, **nineteen** problems at 80 or above, and only twelve at or below 20. Where most figures here are scalpels, he is a general-purpose weapon across the entire applied-statistics half of the board.
**P061 — The hierarchy of hospitals** at 97 is his home problem and reads like a page from his own textbook: partial pooling across 300 hospitals of varying volume, shrinkage proportional to ignorance, and — the part players underrate — propagating uncertainty into the *rankings*, about which he has warned in print that ranks are far noisier than the estimates underneath them. **P125 — The p-value reckoning** also scores 97; he beats Larry Wasserman (92) on it not because Wasserman's technical grasp is weaker but because Gelman's public campaign was more central to the actual historical reckoning. **P212 — Sample from the impossible posterior** (95) is Stan and NUTS. **P296 — Fired by a noisy number** (95) is teacher value-added modeling, where the statistical point — that year-to-year value-added estimates are dominated by noise and that ranking on them is indefensible — is exactly the hierarchical-modeling argument applied to a policy with consequences.
**P137 — The same patients, measured again and again** (93) is longitudinal multilevel modeling; note that Freund scores 6 on this. **P060 — Missing, not at random** (88) is the Rubin inheritance, multiple imputation and missingness mechanisms. **P104 — Are boys more likely than girls?** (88) is a small-effect, large-noise claim of exactly the kind he has spent years dismantling. **P119 — Roll it out in waves** (88) is staged rollout and stepped-wedge design.
His category table explains the breadth: `testing` 61.1, `causality` 60.8, `small-sample` 60.0, `experimental-design` 58.9, `regression` 50.7, `fairness` 63.2. There is essentially no weak spot in applied statistics. `fairness` being that high is worth noting — his applied work on policing data and algorithmic decision-making gives him genuine standing there, not just transferable technique.
The losses are all outside statistics, and they are total. `information` at 7.0 and `systems` at 10.5 are his floor. **P247 — What happened first?** (6) is Lamport's distributed-clock problem; the game's explanation makes the nice observation that "consistency" means something completely different in Bayesian statistics than in distributed systems, and no bridge exists. **P165 — The optimal codebook** (6) is source coding, **P178 — The variables that must be whole** (6) is integer programming, **P036 — Correct the corrupted block** (8) is error-correcting codes, **P160 — Even approximating is hard** (9) is hardness of approximation, and **P046 — Shortest path through the map** (9) is Dijkstra. Combinatorial algorithms, coding theory, and distributed systems are simply not his world.
There is one more weakness a player should anticipate. `high-dim` at 26.0 and `optimization` at 18.0 are low, and they are low for a real reason: Gelman's methodological instincts run toward richly parameterized models of modest-dimensional structured data, not toward the p ≫ n regime where the answer is a sparsity assumption and a convergence rate. On **P018 — p = 20,000, n = 200** you want Wainwright or Candès, not Gelman. And `perception` at 23.0 and `nlp` at 23.0 mean that anything about representations or raw signal belongs to someone else.
The tactical summary is unusually simple. If the problem involves messy real data with group structure, a causal question, an experimental design, or a claim that smells too good, Gelman is very likely your best card. If it involves an algorithm, a code, or a machine, he is close to worthless.