AI History Battle

games

The tournament of strategies

It is 1980 at the University of Michigan, and a political scientist has mailed an invitation to game theorists everywhere: submit a program to play the iterated prisoner's dilemma, round-robin, winner announced publicly. The one-shot game's logic is bleak — defection dominates — yet the repeated game leaves room for something else. Enter the tournament: design a strategy, predict what population of rivals it will face, and explain the result that shocked the field — a four-line program, tit-for-tat, beating every intricate exploiter through being nice, retaliatory, forgiving, and clear. Then draw the real conclusion: under what conditions cooperation among egoists is stable at all. Get it wrong and the mathematics of arms races, trade wars, and evolution keeps predicting only betrayal.

repeated gamesemergence of cooperationsimple beats clever

Who this problem belongs to

The two figures whose methods fit it best, out of 50 in contention.

1928–2015 · midcentury
90

Nash's 1950-51 papers gave game theory the equilibrium concept that explains exactly why the one-shot prisoner's dilemma is bleak: mutual defection is the unique Nash equilibrium even though mutual cooperation pays both players more, because neither can unilaterally improve by deviating from defection. That is the baseline Axelrod's 1980 tournament is built to overturn in the repeated game, where the shadow of future rounds changes which strategies are stable. Nash's own mathematics was for static and repeated games treated abstractly via fixed points, not agent-based tournament simulation, and he did not analyze the specific dynamics of forgiveness, retaliation, and niceness that made tit-for-tat win. He supplied the equilibrium vocabulary the whole problem is stated in; the tournament's empirical, evolutionary answer came from elsewhere.

1916–2001 · midcentury
84

Simon's bounded rationality, developed from the 1940s onward, argued that real agents do not optimize exhaustively but 'satisfice' with simple, robust heuristics under limited computation -- which is precisely the lesson Axelrod's tournament delivered empirically: a four-line rule, tit-for-tat, beat elaborate, computationally heavy strategies submitted by game theorists trying to be clever. Simon's thesis that simplicity and robustness beat exhaustive optimization under uncertainty about what you will face is the tournament's real conclusion, even though he did not run or predict this specific experiment. He worked at Carnegie Mellon on organizational and cognitive decision-making rather than repeated-game tournaments, so the fit is conceptual rather than a direct technical match, but no one in the roster better explains why 'clever' lost.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Game Theory Reinforcement Learning

50 figures are scored on this problem. Draw it in a battle to see where you land.