AI History Battle
Engraved card portrait of Rina Foygel Barber

Rina Foygel Barber

b. 1983 · stat-learning
ask the professor

Conformal prediction; knockoffs; inference after selection

Played by Taylor

1wins
6losses
14.3%win rate

Strongest on

99 Inference after the search 96 When 0.9 must mean ninety percent 96 Prediction intervals without a model 95 Twenty thousand tests at once 90 One test or twenty? 89 p = 20,000, n = 200

Battles

L John Santerre
Counting yeast in the pitching square
L David Silver
Schedule the moonshot
L David Silver
Act on what you cannot see
L David Blei
The parallel text is the teacher
L David Blei
The posterior at web scale
L Pieter Abbeel
Schedule the moonshot
W David Silver
The web of symptoms

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Conformal Prediction Multiple Testing Collaborators

Life and career

Rina Foygel Barber belongs to the generation of statisticians who came of age professionally inside the replication crisis, and her work reads like a sustained answer to it. She took her doctorate in statistics at the University of Chicago, working on high-dimensional and graphical models — her early papers with Mathias Drton on extended information criteria for Gaussian graphical model selection are concerned with a question that already anticipates her later program: when you search over a huge space of models, what do the resulting standard errors and p-values actually mean?

The pivotal move was a postdoctoral position at Stanford with Emmanuel Candès. Candès's group in the early 2010s was the epicenter of high-dimensional statistics with guarantees — compressed sensing, matrix completion, exact recovery under sparsity — and Barber arrived with a question that the sparsity literature had mostly deferred. The lasso will hand you a set of selected variables. Fine. Now what is the error rate of *that set*? Classical inference assumes the hypotheses were fixed in advance; here they were chosen by looking at the data, so every conventional guarantee is void. The answer she and Candès produced, the knockoff filter, was published in 2015 and has since become one of the standard tools of modern selective inference.

She returned to Chicago as faculty, where she is now a chaired professor in the Department of Statistics, and she has become one of the most influential mid-career statisticians in the United States, recognized with the discipline's major early-career honors. Her collaborations — with Candès, with Aaditya Ramdas, with Ryan Tibshirani, with Richard Samworth — form a recognizable school: distribution-free inference, finite-sample validity, and an unusual willingness to publish *negative* results establishing what cannot be done.

That last habit deserves emphasis, because it is rare and it is the thing that most distinguishes her. A great deal of Barber's work consists of impossibility theorems. She has shown that certain forms of conditional predictive inference are unattainable without assumptions; that distribution-free guarantees of particular strengths are provably out of reach in binary regression; that specific desirable properties trade off against each other in ways no cleverer algorithm will escape. This is the mathematically honest version of a discipline under pressure to promise more than it can deliver, and it makes her papers unusually useful: they tell you where the wall is, so you stop walking into it.

Key contributions

**Knockoffs.** The setup: you have a design matrix with many candidate variables, a response, and a selection procedure such as the lasso. You want to control the false discovery rate among the variables you report. The knockoff construction, from the 2015 paper with Candès, builds for each real variable a synthetic "knockoff" copy engineered to have the same covariance structure with the other variables and with each other, but — by construction — no relationship with the response beyond what the original has. The knockoffs are decoys with known null status. Run your selection procedure on the augmented design; every time a knockoff gets selected, you have caught the procedure making a false discovery in a case where you *know* it is false. Counting knockoff selections gives an estimate of the false-discovery count, and a data-dependent threshold on the resulting statistic yields provable FDR control at any target level, in finite samples, without assuming the p-values are valid or independent. The elegance is that it sidesteps post-selection conditioning entirely: rather than analytically characterizing the distribution of a statistic given that it was selected, you supply a control group. The later model-X extension moved the assumptions from the response model to the covariate distribution, making the method applicable when you know a lot about X and little about how Y depends on it — genetics being the motivating case.

**Conformal prediction and distribution-free uncertainty.** Conformal prediction, originating with Vovk, Gammerman and Shafer, produces prediction sets with guaranteed marginal coverage under nothing more than exchangeability of the data — no model assumptions, no asymptotics, valid for any black-box predictor. Barber has been central to its modern development. With Candès, Ramdas and Ryan Tibshirani she introduced the jackknife+ and CV+, which repair leave-one-out prediction intervals so that they carry a rigorous coverage guarantee rather than a heuristic one, at a fraction of the cost of full conformal. With the same collaborators she extended conformal inference *beyond exchangeability*, giving coverage bounds that degrade gracefully — quantified by a measurable distribution-drift term — when the data are not exchangeable, which is the situation in essentially every deployed system. And she proved sharp limits: distribution-free *conditional* coverage, the property everyone actually wants (90% coverage for people like this patient, not merely 90% on average), is impossible without assumptions in continuous settings. Marginal coverage is what you can have for free; conditional coverage must be bought.

**Multiple testing and selective inference more broadly.** Her work on FDR-controlling procedures under dependence, on the interplay between selection and inference, and on the robustness of knockoffs when the covariate distribution is estimated rather than known, together define much of the current state of the art in "I looked at the data before choosing what to test, and I would still like a valid answer."

In battle

Barber has one of the healthiest profiles among the modern statisticians: mean 41.7 across 101 problems, median 40, eleven dominant problems, and only twenty-four at or below 20. She is a specialist, but her specialty — *is this claimed finding real, and how sure can we honestly be?* — turns out to be a question the game asks a great many ways.

Her ceiling is "Inference after the search" (99), which is the knockoff paper as a problem statement: a lasso-selected variable set, a demand for FDR control, no way to condition on the selection event analytically. She built the answer, in the year specified. The other 90s cluster tightly around it. "Prediction intervals without a model" (96) is conformal prediction proper. "When 0.9 must mean ninety percent" (96) is calibration, where finite-sample distribution-free coverage is precisely her guarantee. "Twenty thousand tests at once" (95) and "One test or twenty?" (90) are the multiple-testing core — she scores just under the Benjamini–Hochberg lineage on the origin problems but very high on everything downstream. "p = 20,000, n = 200" (89) is the high-dimensional selection regime her whole toolkit targets. "Fifty examples in the test set" (85) is a lovely fit: with a test set that small, asymptotics are worthless and only finite-sample validity will do — the exact regime conformal methods were built for. And "The bomber that came home" (85) is the survivorship-bias problem; Wald owns the origin, but the modern statement of it is selection bias, and reasoning correctly about what the selection mechanism did to your sample is her home ground.

Her category profile follows cleanly: classification 89.0, high-dimensional 73.5, testing 55.4 across fourteen problems, small-sample 55.2 across sixteen. That small-sample average across sixteen problems is the most valuable number on her sheet — it is a broad, reliable strength rather than a couple of lucky matchups.

Where she loses, she loses on the same axis every 21st-century statistician does: anything that is engineering, computation, or mid-century applied mathematics. "The language for the job" (6) is programming-language design in 1959 — a different layer of the stack and half a century before her field existed. "Will it ever halt?" (9) is computability; her computability average of 9.5 is her worst category. "The parallel text is the teacher" (8) is statistical machine translation, "Best answer before the buzzer" (10) is real-time question answering under a latency budget, and "How many bits must cross the wire" (10) is communication complexity — three problems where the difficulty is building a working system, not certifying a claim. "How much stock to hold" (10) is portfolio construction, a decision-theoretic optimization problem rather than an inferential one, and it points at her real structural weakness: her methods tell you how much you can trust an estimate, not what to *do* given one. Search (11.0), games (11.0), optimization (17.5), RL (17.5) and perception (17.5) are all near the floor for the same reason.

The frank summary is the one her own research culture would endorse: she gives you validity, and validity is not power. Against a problem where a correct parametric model would have delivered a much sharper answer, her assumption-free guarantees look conservative and she can be beaten by someone willing to model. Send her wherever someone has already looked at the data and now wants to make a claim about it.