AI History Battle

fairness

The resume screener learned the past

It is 2018 in Seattle, and an internal audit has killed a machine-learning recruiting tool before it ever officially launched: trained on ten years of the company's own hiring decisions, it learned to downgrade resumes containing the word "women's" and graduates of women's colleges — because it was asked to predict yesterday's decisions, and yesterday was biased. Postmortem the failure completely: why target leakage of historical prejudice is not a bug but the objective faithfully optimized, why scrubbing the obvious tokens fails against paraphrase and proxy, and what evaluation — audited disparate impact against qualified pools, not against past hires — would have caught it pre-deployment. Then answer the harder question: what labeled ground truth for "good hire" could legitimately exist. Get it wrong and automation makes the past permanent.

label biasproxies survive scrubbingwhat is ground truth

Who this problem belongs to

The two figures whose methods fit it best, out of 30 in contention.

b. 1972 · deep-modern
95

O'Neil's Weapons of Math Destruction (2016) essentially predicted this exact Amazon incident before it broke publicly in 2018: a hiring model trained to replicate past decisions faithfully learns and launders whatever bias produced those decisions, and she wrote directly about this case once it surfaced, naming target leakage of historical prejudice as the model working exactly as designed rather than malfunctioning. Her framework for asking who bears the cost, and for insisting that 'the algorithm decided' cannot be the end of an accountability conversation, is precisely the postmortem this scenario demands. Her training as a quant gives her real technical fluency alongside the critique, making her the strongest carrier for this problem's combination of diagnosis and accountability.

b. 1936 · stat-learning
88

Pearl's causal revolution, from Bayesian networks through do-calculus, supplies exactly the missing conceptual tool this postmortem needs: the model was trained to predict a correlational target, past hiring decisions, when what the company actually wanted was a causal quantity, whether a candidate would perform well if hired, and Pearl's framework for distinguishing seeing from doing formalizes precisely why optimizing the former faithfully reproduces the latter's historical bias. His work directly addresses the problem's hardest question, what labeled ground truth for 'good hire' could legitimately exist, since a causal target requires interventional rather than merely observational data. He did not personally audit hiring systems, keeping him just below O'Neil. The throughline from the actual historical record to this exact failure mode is unusually direct for a carrier on this particular list.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Differential Privacy

30 figures are scored on this problem. Draw it in a battle to see where you land.