fairness
The objective is not what you meant
It is a warning first written in 1960 — a founder of cybernetics observing that a machine pursuing a purpose we cannot efficiently revise had better pursue the purpose we actually desire — and it has become the deployment question of the 2020s. A powerful optimizer receives a measurable objective — engagement, cost, test scores — and optimizes exactly that; the gap between metric and intent becomes the system's behavior: clickbait, denial, gaming. Design agents that remain corrigible under objective misspecification: uncertainty over the true objective, deference to human correction, an off-switch the optimizer has no incentive to disable. Prove what the incentives permit. Get it wrong at sufficient capability and the failure is not a bug report — it is a system competently achieving what nobody wanted.
Who this problem belongs to
The two figures whose methods fit it best, out of 37 in contention.
Wiener is not adjacent to this problem, he is its literal author: his 1960 Science essay 'Some Moral and Technical Consequences of Automation' contains almost exactly the warning this scenario opens with, that a machine pursuing a purpose we cannot efficiently revise had better pursue the purpose we actually desire, because 'the penalties for errors of foresight' grow with a system's power and speed. His cybernetics, built around feedback and control loops correcting a system's behavior against a goal, is the direct technical ancestor of the corrigibility question: how do you build a feedback channel a powerful optimizer cannot simply route around. No other carrier on this list originated the warning this problem dramatizes; he is the founder speaking directly to his own prophecy.
Russell's research program on provably beneficial AI is built specifically to answer this problem's exact question: he formalized the case for uncertainty over the true objective, rather than a fixed, possibly misspecified reward, as the mathematical basis for an agent that remains deferential to human correction, and his off-switch game analysis rigorously proves the incentive conditions under which a rational agent chooses not to disable its own shutdown mechanism. His AI: A Modern Approach textbook, co-authored with Norvig, shaped how a generation of researchers think about agent objectives in the first place. This is the single most direct technical match on the entire roster to a scenario asking for corrigibility under objective misspecification, proven incentives to defer, and a description of what capability makes the failure catastrophic rather than merely embarrassing.
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
37 figures are scored on this problem. Draw it in a battle to see where you land.