It is 1990 at IBM's Yorktown labs, and a group of speech researchers is committing linguistic sacrilege: translate French to English with no grammar rules at all, treating translation as a noisy channel — the English sentence garbled into French — and learning everything from millions of sentence pairs of Canadian parliamentary proceedings. Build the statistical translation system: word-alignment models learned from unaligned sentence pairs, a language model to keep the output English, and a decoder to search the space of translations. Then defend the heresy with numbers against linguists who call it barbarism. Get it wrong and translation stays hand-built and brittle for decades; get it right and the data-over-rules argument wins its second great victory — the one language technology never walks back.
Barber's work on conformal prediction and inference after model selection, developed largely in the 2010s, gives rigorous, distribution-agnostic tools for validating a statistical model's predictions, a general methodological contribution with only loose, abstract relevance to validating this problem's alignment-model translations. Her research addresses statistical inference generally rather than translation, word alignment, or the noisy-channel framework specifically, and her career developed decades after this problem's 1990 setting. Her direct engagement with this problem's actual historical moment, its specific technical machinery, or its immediate defense against linguistic orthodoxy is entirely absent, keeping her relevance thin, generic, and largely coincidental to this problem's actual technical content. That gap between modern inference theory and this problem's specific 1990 translation engineering is exactly why her score stays low on this roster.
Blei's topic-modeling work and his development of practical variational inference for large text corpora gave the field statistically rigorous tools for extracting latent structure from language data at scale, a genuine technical descendant of the corpus-driven, data-first instinct this problem's IBM team championed. His methods, like this problem's alignment models, learn structure empirically from large amounts of unannotated text rather than from hand-built linguistic rules. But his career, situated in the 2000s, arrived over a decade after this problem's 1990 setting and addressed document and topic modeling rather than translation or word alignment specifically; his relevance is a strong paradigmatic echo rather than direct engagement with this exact system. That timing gap, arriving over a decade later, is exactly why his score sits below the researchers who actually built this system.
Battle #120 · 8/10/2026, 11:38:52 AM · this result is deterministic: the same two personas on this problem always resolve the same way.