It is 2000, and consumer digital cameras are about to ship with a promise the hardware can barely back: find the faces before the shutter clicks. Detect all faces in family photos — any pose, any lighting, grandmother in shadow, toddler mid-turn — fast enough to run many times a second on a camera's embarrassment of a processor. The design decision is the battle: hand-engineered features composed into a cascade that rejects easy negatives in microseconds, or learned features that need training data and compute the year 2000 does not have in a camera. Justify the choice against the budget. The stakes are the first mass deployment of computer vision into a billion pockets — and every missed dark-skinned face is a bias shipped at consumer scale.
Chose The misspecification audit — right call.
Shalizi's computational mechanics — epsilon-machines, minimal predictive states of stochastic processes — addresses the intrinsic structure of time series and has no purchase on spatial pattern detection; nothing in his research touches images, features, or real-time systems. His genuine relevance is critical: he is among the field's sharpest methodological auditors (his writing on data-mining pathologies, bootstrap honesty, and spurious claims), and the problem's bias clause invites exactly his brand of scrutiny — he would demand stratified error rates and demolish an aggregate-accuracy sales pitch. He also commands enough statistical learning theory to referee the cascade-versus-learned debate intelligently. But refereeing is the whole contribution: no detector, no features, no speed, no data pipeline. A formidable critic with an empty toolbox for this particular job.
Anandkumar's signature results — tensor decompositions with provable guarantees for latent-variable models, and later neural operators for scientific computing — sit far from face detection. The nearest bridges are real but thin: tensor methods generalize the PCA/eigenfaces machinery the era used for appearance modeling, and her NVIDIA-era work on efficient inference speaks abstractly to compute budgets. She also has publicly engaged the demographic-bias failures of face systems, which touches the problem's stakes clause. But her methods assume 2010s optimization and hardware, address parameter recovery rather than sliding-window rejection, and she has no classical-vision, cascade, or embedded record; guaranteed tensor recovery is irrelevant to picking rectangle features that see a toddler mid-turn. Adjacent linear-algebraic machinery and good values, little that runs in the year 2000.
Battle #124 · 8/10/2026, 11:38:59 AM · this result is deterministic: the same two personas on this problem always resolve the same way.