AI History Battle

perception

A thousand categories, one bet

It is 2012 in Toronto, and the ImageNet challenge — a million images, a thousand categories — has become computer vision's annual referendum, ruled by hand-engineered features whose error rate improves a fraction of a point a year. The bet on your desk is unfashionable: a deep convolutional network, tens of millions of parameters, trained end-to-end on two consumer gaming GPUs, with tricks — rectified units, dropout, aggressive augmentation — to keep a model this large from memorizing. The field's establishment considers this lineage refuted. Halve the error rate, publicly. Get it wrong and neural networks absorb one more decade of dismissal; get it right and every vision pipeline on earth is rebuilt within thirty-six months — the discontinuity by which the deep-learning era dates itself.

end-to-end learningGPU scalebenchmark discontinuity

Who this problem belongs to

The two figures whose methods fit it best, out of 53 in contention.

b. 1947 · deep-modern
98

Hinton is not analogous to this problem, he is its principal author: with graduate students Alex Krizhevsky and Ilya Sutskever at Toronto, he built AlexNet, the deep convolutional network that halved ImageNet's error rate in 2012 and ended computer vision's decade of hand-engineered-feature dominance almost overnight. His decades of unfashionable advocacy for neural networks through the field's coldest years, backpropagation popularization, Boltzmann machines, made him the one senior researcher positioned to bet on this architecture publicly when the establishment considered the lineage refuted. The specific tricks that kept a model this large from memorizing, dropout regularization among them, were developed in his own lab in the years just before. Training end-to-end on two consumer GPUs was the group's own engineering choice. No carrier in this pool owns this exact result more directly.

b. 1986 · deep-modern
95

Sutskever was the hands-on engineer who made AlexNet actually train: as Hinton's graduate student, he wrote much of the CUDA implementation that let a network this large, tens of millions of parameters, run end-to-end on two consumer gaming GPUs rather than the specialized supercomputers vision researchers assumed were required. His technical judgment about rectified linear units accelerating convergence and his broader instinct, later crystallized as the scaling hypothesis, that flexible architectures trained on enough data and compute beat hand-engineering, is this problem's entire thesis stated as a research philosophy. He co-authored the 2012 NeurIPS paper with Krizhevsky and Hinton that is this problem's exact historical event, making him one of the two most directly responsible carriers in this entire pool for this specific result.

Fought here

Larry Wasserman beat Noam Chomsky 18–15

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Computer Vision Neural Networks ImageNet

53 figures are scored on this problem. Draw it in a battle to see where you land.