AI History Battle

causality

Thirty percent chance of rain

It is 1978 at the National Weather Service, and an underappreciated discipline has quietly emerged: forecasters have issued probability-of-precipitation numbers for a decade, and someone must finally ask whether the numbers mean anything. Audit them: when the forecast says seventy percent, does it rain seven times in ten? Build the calibration analysis, score the forecasts with a rule that rewards honesty — where hedging toward fifty percent costs points and both overconfidence and timidity are punished — and separate calibration from resolution, since a forecaster who always recites the climatological base rate is calibrated and useless. Get it wrong and probability numbers become decoration; get it right and weather forecasting becomes the rare institution whose stated uncertainties are true — the existence proof every other forecasting field needs.

calibrationproper scoring rulesinstitutional verification

Who this problem belongs to

The two figures whose methods fit it best, out of 61 in contention.

1915–2000 · early-stat
92

Tukey chaired the panel work that helped push the U.S. Weather Bureau toward issuing numerical probability-of-precipitation forecasts in the 1960s, arguing that a forecaster's honest uncertainty is itself information the public deserves rather than a hedge to be laundered into "cloudy." His statistical instinct for confronting a method with its own errors, exploratory data analysis built on residuals and diagnostic plots, is exactly the discipline this problem needs to audit a decade of PoP numbers against actual outcomes. He also understood, decades before anyone formalized proper scoring rules for meteorology, that a scoring system must punish both overconfidence and cowardly hedging toward the climatological mean, or it rewards the wrong behavior. Few careers connect this directly to institutionalized probabilistic forecasting.

1919–2010 · stat-learning
85

Blackwell's game-theoretic work on comparison of statistical experiments and, later, the Foster-Vohra proof that calibration is achievable via Blackwell's own approachability theorem, is the deepest theoretical bedrock underneath this problem: a forecaster facing an adversarial or indifferent nature can always steer toward well-calibrated predictions using strategies descended directly from his 1956 approachability result. His decision-theoretic sensibility, treating a forecast as an action scored by a loss function rather than a belief stated for its own sake, is precisely the register a proper scoring rule requires. He did not personally audit the National Weather Service, and his approachability work predates this specific 1978 application by two decades, so the transfer is foundational rather than applied.

Fought here

Larry Wasserman beat John Hopfield 64–18

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Causal Inference Bayesian Networks

61 figures are scored on this problem. Draw it in a battle to see where you land.