experimental-design
When you can't randomize
It is a moment every policy analyst dreads: a program has already been rolled out, city by city, chosen for political convenience rather than by any coin you controlled — and now you must say whether it worked. Randomization, the one clean guarantee, was never on the table. Estimate the program's causal effect anyway, and state with total clarity exactly which assumptions are doing the work that randomization normally would: what you must believe about how cities were selected, which confounders you must rule out. The assumptions ARE the analysis. Get them wrong and a program that did nothing gets scaled nationwide on your say-so, or a program that worked gets killed — and no one will see the flawed assumption buried under the confident estimate.
Who this problem belongs to
The two figures whose methods fit it best, out of 34 in contention.
This problem is essentially Rubin's life work stated as a scenario. From the mid-1970s onward he built the potential-outcomes framework precisely for causal questions where randomization was absent: define the counterfactual estimand first, then state the assignment mechanism explicitly. His 1983 work with Rosenbaum on the propensity score gave the standard tool for observational program evaluation — model how cities were selected into treatment, balance on that score, and check overlap before estimating anything. Ignorability, the assumption 'doing the work randomization normally would,' is his vocabulary; so are sensitivity analyses that ask how strong a hidden confounder must be to overturn the estimate. He spent decades applying exactly this to policy and legal disputes. The problem's demand that the assumptions BE the analysis is the Rubin causal model's founding principle.
Pearl arrives from a different direction — computer science and Bayesian networks in the 1980s — but by the 1990s his do-calculus made the assumptions in this problem literally drawable. Where Rubin writes ignorability as a conditional-independence statement, Pearl draws the causal graph of how cities were selected, marks the confounders, and applies back-door and front-door criteria to determine whether the program effect is identifiable at all from observational data — and if not, exactly which unmeasured arrow blocks it. That is the problem's core demand: total clarity about which assumptions substitute for randomization. His framework also cleanly separates identification from estimation, so the flawed-assumption-buried-under-a-confident-estimate failure mode is exactly what the graph exposes. He is marginally behind Rubin only because propensity-score practice dominates applied city-level program evaluation.
Fought here
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
34 figures are scored on this problem. Draw it in a battle to see where you land.