AI History Battle

fairness

The objective is not what you meant

It is a warning first written in 1960 — a founder of cybernetics observing that a machine pursuing a purpose we cannot efficiently revise had better pursue the purpose we actually desire — and it has become the deployment question of the 2020s. A powerful optimizer receives a measurable objective — engagement, cost, test scores — and optimizes exactly that; the gap between metric and intent becomes the system's behavior: clickbait, denial, gaming. Design agents that remain corrigible under objective misspecification: uncertainty over the true objective, deference to human correction, an off-switch the optimizer has no incentive to disable. Prove what the incentives permit. Get it wrong at sufficient capability and the failure is not a bug report — it is a system competently achieving what nobody wanted.

objective misspecificationcorrigibilityincentives to defer

Who this problem belongs to

The two figures whose methods fit it best, out of 37 in contention.

1894–1964 · midcentury
99

Wiener is not adjacent to this problem, he is its literal author: his 1960 Science essay 'Some Moral and Technical Consequences of Automation' contains almost exactly the warning this scenario opens with, that a machine pursuing a purpose we cannot efficiently revise had better pursue the purpose we actually desire, because 'the penalties for errors of foresight' grow with a system's power and speed. His cybernetics, built around feedback and control loops correcting a system's behavior against a goal, is the direct technical ancestor of the corrigibility question: how do you build a feedback channel a powerful optimizer cannot simply route around. No other carrier on this list originated the warning this problem dramatizes; he is the founder speaking directly to his own prophecy.

b. 1962 · ai-classic
97

Russell's research program on provably beneficial AI is built specifically to answer this problem's exact question: he formalized the case for uncertainty over the true objective, rather than a fixed, possibly misspecified reward, as the mathematical basis for an agent that remains deferential to human correction, and his off-switch game analysis rigorously proves the incentive conditions under which a rational agent chooses not to disable its own shutdown mechanism. His AI: A Modern Approach textbook, co-authored with Norvig, shaped how a generation of researchers think about agent objectives in the first place. This is the single most direct technical match on the entire roster to a scenario asking for corrigibility under objective misspecification, proven incentives to defer, and a description of what capability makes the failure catastrophic rather than merely embarrassing.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Differential Privacy

37 figures are scored on this problem. Draw it in a battle to see where you land.