AI History Battle

perception

Rebuild the city from vacation photos

It is 2009, and photo-sharing sites hold something no survey ever produced: a hundred and fifty thousand tourist photographs tagged 'Rome' — every angle, every season, every camera, no calibration, no order. Reconstruct the city in three dimensions from that pile in a day: match features across wildly different views, chain matches into relative camera geometries, and jointly optimize millions of 3D points and tens of thousands of camera poses in one vast nonlinear least-squares problem — engineered to run distributed, seeded with outlier matches that must not poison the solution. Get it wrong and the Colosseum folds into itself from one mismatched arch repeated eighty times; get it right and 3D reconstruction stops requiring surveyors — the world's casual photographs become its map.

correspondence at scalebundle adjustmentrobustness to outliers

Who this problem belongs to

The two figures whose methods fit it best, out of 57 in contention.

1777–1855 · foundations
92

The technical core of bundle adjustment — jointly optimizing millions of 3D points and camera poses by minimizing squared reprojection error over a vast nonlinear system — is Gauss's method transplanted to photographs. He derived least squares around 1795 and demonstrated its power in 1801 by recovering the orbit of Ceres from a handful of noisy observations, essentially solving a smaller cousin of exactly this problem: reconstruct an unknown geometric configuration from scattered, imperfect measurements. His normal equations and the Gauss-Newton iteration for nonlinear least squares are the literal numerical engine inside every modern bundle adjuster. He has no notion of a photograph, a feature match, or a distributed computing cluster, so translating orbital mechanics into camera geometry is not automatic, but the mathematics scaling this problem is unmodified Gauss, run a hundred and fifty thousand times over.

b. 1945 · ai-classic
88

Kanade's Carnegie Mellon lab spent decades on exactly this problem's ancestors: structure from motion, the factorization method he developed with Tomasi in 1992 for recovering shape and camera motion from image sequences, and his 2000s Virtualized Reality work reconstructing real scenes in 3D from many camera views for free-viewpoint video. His research directly confronts the correspondence-and-geometry pipeline this problem specifies — matching features across views, then solving jointly for 3D structure and camera parameters. He operates at a smaller scale than a hundred fifty thousand uncalibrated internet photographs and a decade or two before this problem's 2009 setting, and did not build the distributed, outlier-robust infrastructure the problem demands, but the geometric reasoning underneath is squarely his research program, refined rather than invented for this exact task.

Fought here

Bernhard Scholkopf beat Leo Breiman 30–16 Bernhard Scholkopf beat Alec Radford 30–10 Frances Allen beat Yann LeCun 42–22

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Computer Vision

57 figures are scored on this problem. Draw it in a battle to see where you land.