nlp
A recognizer for a language of ten speakers
It is 2019, and the language technologies reshaping the world serve perhaps a hundred of humanity's seven thousand languages. A community of ten thousand speakers wants what English has — transcription, translation, search — but there is no treebank, no parallel corpus, no web crawl to feed the data-hungry methods; there are a few hours of recorded elders and a dictionary compiled by a missionary in 1930. Build language technology under this budget: transfer from related and unrelated high-resource languages, exploit phonetic universals, learn from the community's own corrections, and measure honestly on the language itself rather than a proxy. Get it wrong and the digital age becomes another engine of language death — the tools that could document a language arriving only after it is gone.
Who this problem belongs to
The two figures whose methods fit it best, out of 47 in contention.
Reddy's career is unusually well matched to this exact problem: his pioneering work on continuous speech recognition at Carnegie Mellon from the 1970s onward established the statistical and acoustic modeling foundations any language technology needs, and his later career was substantially devoted to bringing computing and language technology to underserved populations globally, including work on affordable technology access and the Universal Digital Library aimed explicitly at communities the mainstream technology industry ignores. His Turing Award-winning insistence that language systems be evaluated on real-world performance for real communities, not benchmark elegance, matches this problem's demand to measure success on the language itself rather than a proxy. Reddy worked primarily on well-resourced languages technically, even as his advocacy targeted global equity, so the specific transfer-learning techniques this problem needs are not entirely his own toolkit.
Gebru's founding of the Distributed AI Research Institute after leaving Google, and her sustained advocacy for language technology that serves communities outside the small set of languages the industry prioritizes, engages this problem's 'whose languages' axis more directly than any other figure on this roster. Her work on datasheets for datasets and model cards insists that any language technology honestly document what data it was built from and whose interests it serves, directly relevant to this problem's demand to measure success on the community's own terms rather than a proxy benchmark built for high-resource languages. Gebru's own technical background is in computer vision and later AI ethics and accountability rather than low-resource speech and translation engineering specifically, so the acoustic and transfer-learning techniques this problem needs are not her direct contribution.
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
47 figures are scored on this problem. Draw it in a battle to see where you land.