AI History Battle

nlp

What is in the training data?

It is 2021, and language models are trained on crawls so large that no one — including their builders — can say what is in them: which voices are overrepresented and which filtered out by "quality" heuristics that correlate with dialect and class, how much toxicity is baked in, whose personal text was swept up. The models are shipping anyway. Do the curation science: characterize a web-scale corpus statistically, audit what the filtering removed and from whom, document the corpus so downstream users inherit knowledge instead of mystery, and weigh the honest question of whether ever-larger undocumented crawls are a research direction or a liability. Get it wrong and every model trained downstream launders the corpus's omissions into authoritative-sounding output — at the scale of everyone.

corpus curationdocumentationscale vs accountability

Who this problem belongs to

The two figures whose methods fit it best, out of 55 in contention.

b. 1983 · deep-modern
98

Gebru's 2018 'Datasheets for Datasets' proposal and her 2021 'Stochastic Parrots' paper, co-authored while she was pushed out of Google over exactly this dispute, are not analogous to this problem, they are its direct origin. Datasheets ask precisely what this scenario demands: provenance, collection process, and known gaps documented before a corpus ships, so downstream users inherit knowledge instead of mystery. Stochastic Parrots specifically names the risk of undocumented web-scale crawls encoding toxicity and skewed demographics under the cover of 'quality' filters that correlate with dialect and class. She subsequently founded DAIR to institutionalize exactly this kind of independent audit work outside corporate incentives. No other career on this card maps this cleanly onto the problem's own language, which is why she carries the maximum score with no real competitor.

b. 1972 · deep-modern
88

O'Neil's 'Weapons of Math Destruction' (2016) built the general vocabulary this problem needs: opaque, large-scale models trained on data nobody fully audits, whose errors compound invisibly against the already disadvantaged. Her career pivot from quantitative finance to public-interest algorithmic auditing models exactly the posture the problem demands, treating a model's training corpus as a claim requiring evidence rather than an asset to ship quietly. She built practical auditing frameworks (O'Neil Risk Consulting & Algorithmic Auditing) for exactly this kind of accountability work across industries. Her focus is broader than NLP corpora specifically, spanning credit, hiring, and policing data, so her score reflects outstanding general accountability method transferred into this domain rather than hands-on web-corpus linguistics, the reason she sits just below Gebru's direct authorship.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Language Models

55 figures are scored on this problem. Draw it in a battle to see where you land.