AI History Battle

perception

Replace the acoustic model

It is 2012, and speech recognition has spent twenty years in a refined stasis: Gaussian mixtures over hand-designed acoustic features, decoded through hidden Markov models, improving by fractions as an industry tunes the same machinery. The proposal is to tear out the acoustic model and drop in a deep network trained on raw-ish spectral input — heresy from a community the speech establishment considers twice-buried, and it must prove itself on the industrial benchmarks, under the same decoders, at deployable speed. Show the word-error rate falling by a third, replicated across research groups and languages, not cherry-picked. Get it wrong and voice interfaces stay a novelty that mistranscribes anyone with an accent; get it right and within five years dictation, assistants, and captioning cross the threshold of everyday use.

hybrid systemslearned features vs engineeredindustrial benchmarks

Who this problem belongs to

The two figures whose methods fit it best, out of 57 in contention.

b. 1947 · deep-modern
98

This problem is Hinton's own 2012 desk. He was a lead author, with George Dahl, Abdel-rahman Mohamed, and colleagues at Toronto and Microsoft, IBM, and Google research groups, of the survey paper "Deep Neural Networks for Acoustic Modeling in Speech Recognition," which documented exactly this proposal: replace Gaussian mixture acoustic models with deep networks trained on spectral features, decoded through the same hidden Markov model backbone, and show word-error-rate gains replicated across research groups and languages, not cherry-picked. His pretraining tricks using restricted Boltzmann machines made deep networks trainable before this deployment, and the twenty-year GMM stasis is precisely the "twice-buried" establishment he had argued against for years. He proposed, built, and validated this exact system; only the honest note that a large team executed the replication keeps this from a bare 100.

b. 1968 · systems
90

Dean is a co-author on the same 2012 "Deep Neural Networks for Acoustic Modeling" survey that documents this problem's exact proposal, and his infrastructure work at Google — the systems that would let deep acoustic models run at the speed and scale industrial decoders require — is exactly the "deployable speed" and "industrial benchmarks" half of this problem's demand. His large-scale distributed training systems, developed in this same period, made training deep networks on production-scale speech data computationally tractable rather than a laboratory curiosity. He is not the acoustic-modeling researcher who proposed replacing Gaussian mixtures with deep networks — that is squarely Hinton's and his collaborators' contribution — so the systems half of this problem's success belongs to Dean while the modeling insight belongs to others, which is the honest reason he sits just below Hinton.

Fought here

Carlos Guestrin beat John Santerre 20–0

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Hidden Markov Models

57 figures are scored on this problem. Draw it in a battle to see where you land.