AI History Battle

It is 2009, and photo-sharing sites hold something no survey ever produced: a hundred and fifty thousand tourist photographs tagged 'Rome' — every angle, every season, every camera, no calibration, no order. Reconstruct the city in three dimensions from that pile in a day: match features across wildly different views, chain matches into relative camera geometries, and jointly optimize millions of 3D points and tens of thousands of camera poses in one vast nonlinear least-squares problem — engineered to run distributed, seeded with outlier matches that must not poison the solution. Get it wrong and the Colosseum folds into itself from one mismatched arch repeated eighty times; get it right and 3D reconstruction stops requiring surveyors — the world's casual photographs become its map.

correspondence at scalebundle adjustmentrobustness to outliers
b. 1968
tapped · ask the professor
30

Scholkopf's systematization of the kernel trick and kernel methods more broadly gives him deep comfort with high-dimensional feature spaces and robust statistical estimation, which offers some abstract conceptual purchase on this problem's correspondence-matching and outlier-robustness demands. His later work on causal machine learning shows a broader interest in recovering structure from data under uncertainty, and his research at the Max Planck Institute for Intelligent Systems touched robotics and vision applications adjacent to this kind of geometric estimation, though not this specific pipeline. He never worked on structure from motion, camera geometry, or large-scale distributed bundle adjustment specifically, and kernel methods are not the actual tool used to solve this problem's geometric optimization, so the connection remains at the level of general statistical sophistication rather than directly applicable published research.

b. 1986
was tapped · ask the professor
10

Radford's work on GPT-style generative pretraining and CLIP's joint image-text representation learning centers on large-scale self-supervised learning from massive datasets, a philosophy built for learned mappings rather than this problem's classical, unsupervised geometric reconstruction from photographic correspondences. There is no large training corpus or contrastive objective involved in jointly solving for camera poses and 3D points from matched features; the 2009 pipeline's constraints are also distant from his GPU-scale context. CLIP's contact with images gives a thin domain adjacency, but nothing in his published work addresses structure from motion directly. CLIP's contrastive matching of images to text does share a family resemblance with feature correspondence across images, since both learn what counts as the same underlying thing seen differently, even though the training objective, scale, and era are entirely different from this problem's classical pipeline.

Head to head 20 over 2 battles
Read Scholkopf Read Radford Leaderboard

Battle #85 · 8/10/2026, 11:36:48 AM · this result is deterministic: the same two personas on this problem always resolve the same way.