It is the same labels-are-expensive world, but now you hold the pen: you may choose which thirty of the thirty thousand points get sent to the expert for labeling, and you may choose them adaptively, each new label informing the next request. A point deep inside a cluster teaches you little; a point near the uncertain boundary can be worth dozens of random ones. Design the querying strategy and prove it beats labeling thirty points at random. The guarantee must be provable, not merely plausible. Get it wrong and you spend your tiny, precious label budget on redundant points and learn no faster than blind sampling — when every label costs an expert's hour, choosing the right questions is the entire game, and 'provably better than random' is the bar.
Chose The size principle — right call.
Tenenbaum's science is about exactly this problem's astonishment: how learners generalize from absurdly few labels. His Bayesian concept-learning work (the number game, word learning) formalizes how strong priors plus the size principle extract maximum evidence from tiny samples, and his lab's rational-analysis tradition includes optimal experiment design — choosing the query that maximizes expected information gain — as a model of human question-asking, tested against behavior. He also co-authored foundational manifold work (Isomap, 2000), so the unlabeled cloud's geometry is literally his toolkit. The gaps: his guarantees are Bayesian-rational, model-internal claims, not distribution-free proofs against random sampling; his systems are cognitive models, not budget-constrained annotation pipelines; and the PAC-style bar the problem sets belongs to a different literature. High conceptual fit, incomplete on the required theorem.
Pearson founded mathematical statistics on the large-sample, fixed-data paradigm: gather ample observations, then characterize them with correlation, chi-squared goodness-of-fit, and the method of moments. Every element of this problem inverts his world. His asymptotics collapse at thirty labels; his framework has no concept of a hypothesis class, a decision boundary, or generalization error; and — most fatally — his statistics is entirely non-adaptive: data arrive, they are not chosen, and the idea of sequential design responsive to interim results belongs to Fisher's experimental-design school and Wald's sequential analysis, both after him and neither his. The thirty thousand unlabeled points would strike him as material for biometric curve-fitting, not as geometry constraining a classifier. He is a floor case here: foundational giant, toolkit almost perfectly orthogonal to the task.
Battle #76 · 8/10/2026, 11:36:22 AM · this result is deterministic: the same two personas on this problem always resolve the same way.