AI History Battle
XOR classification

It is 1969, and four points in the plane — two labeled one class, two the other, arranged as exclusive-or — are about to reshape the funding of an entire field. No straight line separates them; the single-layer perceptron, the great hope of machine intelligence, provably cannot solve this toy. Solve it — by transforming the representation until the classes come apart — and explain what the public demonstration of this failure did: it helped freeze neural-network research for the better part of two decades. The stakes here are historical, not just technical. Get the lesson wrong and you either overclaim for a method that has a hard wall, or abandon a whole paradigm over a limitation that richer representations dissolve. Four points decided who got money for twenty years.

nonlinearpredictrepresentation
b. 1979
tapped · ask the professor
18

Clauset's toolkit — rigorous maximum-likelihood fitting of power laws, community detection, and statistical inference on networks — matured in the 2000s and aims at a different target entirely: structure in large relational data, not the representational limits of learning machines. Nothing in his methods addresses why a linear threshold unit cannot compute parity, nor how to construct the hidden features that make four points separable. His honest-hypothesis-testing ethos is a distant cousin of the XOR lesson (do not overclaim what your model class can express), and a modern network scientist could of course recite the Minsky-Papert story from any textbook. But reciting history is not the same as owning the methods. His statistical machinery gains no traction on a four-point boolean toy; this problem lives outside his field's borders.

b. 1986
was tapped
44

Vaswani's transformer (2017) is stacked nonlinear representation learning at industrial scale — every feed-forward block inside it solves XOR-class problems incidentally, thousands of times over, and the whole architecture is a monument to the principle that learned representations dissolve linear-separability walls. So the lesson of 1969 is fully absorbed into his working assumptions. But that is inheritance, not method-fit: pointing a transformer at four boolean points is absurd overkill that illuminates nothing, and his specific contributions — attention, positional encoding, large-scale sequence modeling — address long-range dependency in language, not minimal counterexamples in learnability. He was born after the perceptron controversy ended and engages its history as textbook background. Grade him as a beneficiary of the resolution who could demonstrate it trivially, but whose actual tools are tuned to a different problem three eras later.

Head to head 52 over 7 battles
Read Clauset Read Vaswani Leaderboard

Battle #128 · 8/10/2026, 11:39:15 AM · this result is deterministic: the same two personas on this problem always resolve the same way.