AI History Battle

high-dim

The interpolator that should have failed

It is 2019, and deep learning has placed an awkward exhibit on statistical learning theory's doorstep: networks with far more parameters than data points, trained to exactly zero error — interpolating even corrupted labels — that nonetheless generalize well, while the textbook bias-variance story predicts disaster past the interpolation point. Worse, test error can descend a second time as models grow beyond it. Resolve the anomaly honestly: what the classical uniform-convergence bounds actually claimed, which implicit regularization the optimizer contributes uninvited, in what regimes interpolation is provably benign, and where the folklore version of the new story overreaches in turn. Get it wrong twice — theory dismissing the phenomenon, or practice concluding overfitting was never real — and the field's guidance to practitioners becomes superstition in both directions.

interpolation regimeimplicit regularizationtheory meets anomaly

Who this problem belongs to

The two figures whose methods fit it best, out of 63 in contention.

b. 1966 · stat-learning
96

Bartlett is a direct author of the technical resolution this problem asks for: his 2020 paper with Long, Lugosi, and Tsigler, 'Benign Overfitting in Linear Regression,' proves precisely when an interpolating estimator, one achieving zero training error including on noisy points, can still generalize well, identifying the spectral conditions on the data's covariance structure that separate benign from catastrophic interpolation. His decades of statistical learning theory work on margins and generalization bounds also explain what classical uniform-convergence theory actually claimed and where it silently assumed non-interpolation, addressing this problem's demand to state honestly what the old theory promised. His Berkeley-era standing placed him at the center of the exact 2019 debate this problem describes. His score reflects primary, technically precise authorship of the phenomenon's resolution, not adjacent relevance.

b. 1974 · stat-learning
92

Srebro's research on implicit regularization, how gradient descent on overparameterized models is biased toward simple, well-generalizing solutions even without explicit penalty terms, is exactly the 'implicit regularization the optimizer contributes uninvited' this problem demands be identified. His work characterizing the margin-maximizing bias of gradient descent on separable data, and his broader program on norms and generalization in overparameterized learning, directly explains why networks trained to interpolation can still generalize: not despite zero training error but because of which particular zero-training-error solution the optimizer finds. His papers from exactly the 2018-2019 period engage this problem's anomaly as live research. His score reflects primary authorship of the implicit-regularization half of this problem's required resolution.

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Statistical Learning Theory Regularization Deep Learning

63 figures are scored on this problem. Draw it in a battle to see where you land.