AI History Battle

nlp

The topics in the archive

It is 2003, and libraries are digitizing faster than anyone can read: a century of newspaper archives, millions of articles, and historians who can sample a shelf but never survey the corpus. Discover the themes in the archive and their rise and fall over decades — unsupervised, because no one can label a century, and interpretable, because the end user is a historian who must be able to look at a theme and name it. The model must be honest probability, not clustering folklore: documents as mixtures of topics, topics as distributions over words, uncertainty carried rather than hidden. The stakes are a new instrument for the humanities — or, done badly, a machine for generating confident narratives about the past that are artifacts of the fit.

latent structureunsupervised interpretability

Who this problem belongs to

The two figures whose methods fit it best, out of 50 in contention.

b. 1976 · stat-learning
98

Blei, with Ng and Jordan, published latent Dirichlet allocation in 2003, the exact paper this problem describes: documents as mixtures of topics, topics as distributions over words, fit via variational inference to remain computationally tractable at corpus scale while carrying honest uncertainty rather than hiding it behind hard clustering assignments. This is not an analogous or adjacent contribution; it is the literal technical object and publication year the problem specifies, addressing precisely the interpretability-for-historians requirement LDA was explicitly designed to satisfy better than earlier clustering approaches. No other carrier in this batch can claim to have built the actual mathematical model and inference procedure this problem calls for, at the exact historical moment specified.

b. 1956 · stat-learning
88

Jordan's decades-long research program on graphical models and variational inference, including being Blei's doctoral advisor and a co-author on the original LDA paper itself, places him at the direct technical and institutional center of this exact problem. He systematized the mean-field variational methods that make LDA computationally tractable at corpus scale and provided much of the theoretical language, exponential families, plate notation, approximate inference, that makes the model coherent as a principled Bayesian object rather than an ad hoc heuristic. His fingerprints are on the actual paper this problem describes, making him one of the strongest possible carriers for this exact 2003 unsupervised topic-modeling problem. Few carriers in this entire roster have a more direct documented connection to the specific paper this problem describes.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Distributions

50 figures are scored on this problem. Draw it in a battle to see where you land.