AI History Battle

fairness

Fired by a noisy number

It is 2011 in Washington D.C., and the school district is dismissing teachers on the strength of value-added scores — statistical estimates of each teacher's contribution to student test growth, computed from a year or two of small, non-randomly-assigned classes. The scores bounce: a teacher stellar one year lands in the bottom quintile the next, and the confidence intervals, when anyone computes them, span most of the scale. Audit the instrument before it is a weapon: quantify the year-to-year reliability, the shrinkage that honest small-sample estimation demands, the sorting bias from non-random assignment, and the Goodhart dynamics once pay and employment bend the metric. Then say what stakes this measurement can bear. Get it wrong and careers end on noise — while teaching to the test becomes the surviving skill.

reliability of rankingsshrinkageGoodhart under stakes

Who this problem belongs to

The two figures whose methods fit it best, out of 30 in contention.

b. 1965 · stat-learning
95

Gelman's hierarchical Bayesian modeling, developed and popularized through decades of applied work and the Stan software he co-created, is precisely the statistical corrective this scenario needs: value-added scores computed from one or two small, non-randomly-assigned classes are exactly the small-sample, high-variance estimation problem hierarchical models exist to fix, borrowing strength across teachers to shrink noisy individual estimates toward a more reliable common signal rather than trusting each raw score in isolation. He has written directly and publicly about the statistical illiteracy of exactly this kind of high-stakes ranking from noisy year-to-year data, making him the strongest possible carrier for both the technical shrinkage fix and the public communication of why the unshrunk scores were dangerous to act on.

b. 1938 · stat-learning
88

Efron's empirical-Bayes methods, developed from the 1970s onward and building on Stein's earlier shrinkage insight, are the direct statistical machinery this problem calls for: estimating each teacher's true contribution by borrowing strength from the entire distribution of teachers rather than trusting a single small-sample estimate, exactly the honest shrinkage this scenario says was never applied before careers ended. His bootstrap also supplies a principled way to quantify how wide the confidence intervals on a one-year value-added score actually are, which the scenario notes 'span most of the scale' once anyone bothers to compute them. He did not personally study teacher evaluation, keeping him just below Gelman. The throughline from the actual historical record to this exact failure mode is unusually direct for a carrier on this particular list.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Confidence Intervals

30 figures are scored on this problem. Draw it in a battle to see where you land.