AI History Battle

causality

Concepts from three examples

It is 2011 at MIT, and the gap is embarrassing: state-of-the-art learning systems need thousands of labeled examples to recognize a category, while a three-year-old down the hall at the daycare sees three examples of a new word — "that's a wug, there's another wug" — and generalizes correctly, immediately, including to cases she has never seen. Model this one-shot concept learning as Bayesian program induction: structured priors over how concepts are built, likelihoods that explain why three examples suffice, inference that lands where the child lands. The stakes run in both directions — a computational account of human concept learning for cognitive science, and an existence proof that data-hungry statistics is not the only path, which the deep-learning era badly needs to hear.

structured priorscognitive

Who this problem belongs to

The two figures whose methods fit it best, out of 59 in contention.

b. 1972 · deep-modern
97

This is Tenenbaum's problem in the most literal sense: it is set in his lab. His 1999 thesis derived the size principle explaining why a few examples produce confident graded generalization; his 2001 work with Griffiths unified generalization models under Bayes; Kemp-Perfors-Tenenbaum built hierarchical priors that learn to learn; and by 2011 he and Brenden Lake were assembling the Bayesian program-induction account of one-shot character learning that became the 2015 Science paper — structured priors over how concepts compose, likelihoods that make three examples sufficient, inference validated against human judgments. The 2011 Science review with Kemp, Griffiths, and Goodman states this problem's manifesto verbatim: how does the mind get so much from so little? The only headroom above him is that the program remains incomplete — inference is expensive, scope contested. On this turf, he is the turf.

1903–1987 · early-stat
86

Kolmogorov owns both pillars this problem stands on. His 1933 axiomatization makes the Bayesian calculus rigorous, and — decisively — his 1960s algorithmic complexity theory defines the prior the problem needs: the probability of a concept as a function of its shortest description. 'Structured priors over how concepts are built' is Solomonoff-Kolmogorov induction wearing cognitive clothing, and the size principle explaining why three examples suffice is a likelihood argument he had the tools to state exactly. He thought about description complexity to the end of his life. The era gaps are real: Kolmogorov complexity is uncomputable, so the 2011 work's approximate inference over program traces and its behavioral validation against children's judgments lie beyond his methods. But the theoretical spine of program induction is his.

Fought here

Jeff Dean beat Partha Niyogi 53–11 Timnit Gebru beat John Santerre 16–0

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Causal Inference Bayesian Networks

59 figures are scored on this problem. Draw it in a battle to see where you land.