It is 2017, and the uncomfortable laboratory result has walked outdoors: image classifiers that surpass humans on benchmarks can be inverted, their gradients yielding perturbations invisible to people that flip any prediction — and now a few printed stickers on a physical stop sign make a driving-grade detector read SPEED LIMIT 45 from every angle and distance a camera will pass through. Explain the vulnerability — high-dimensional linearity, not exotic overfitting — then attack and defend honestly: craft the physical, robust perturbation, and evaluate defenses against an adversary who knows them, because obscured gradients have already embarrassed a round of published fixes. Get it wrong and perception models certified on clean test sets ship into safety-critical roads while their failure modes are printable at home.
Abbeel's research on robot learning and deep reinforcement learning for physical manipulation depends on exactly the kind of robust, safety-critical perception this problem's stop-sign attack threatens, and his practical robotics instinct that a corrupted perception estimate becomes a physically dangerous action aligns directly with this problem's stakes. His Berkeley graduate work on state estimation from noisy sensors gives him some hands-on experience with robustness under real-world uncertainty. He was not part of the specific 2017 adversarial- examples and physical-attack research program this problem describes, and his own contributions are in learned control policies rather than the gradient-based attack-and-defense engineering this problem requires, leaving him a downstream stakeholder rather than a direct contributor.
Urtasun's research and her company Waabi are built on exactly the application this problem concerns: perception systems for autonomous vehicles that must be certified safe under real, adversarial-quality conditions rather than clean benchmark data, since a misread stop sign in her domain is not an academic embarrassment but a potential collision. Her professional instinct to treat a perception failure as a safety-critical event rather than a statistical curiosity matches this problem's stakes precisely. She was not part of the specific 2017 adversarial-examples research program that discovered and formalized this vulnerability, and her own published contributions center on general robust perception rather than the gradient-based attack-and- defense methodology this problem requires, so the concrete adversarial-machine-learning deliverable belongs to security-focused researchers rather than to her directly.
Battle #136 · 8/10/2026, 11:39:51 AM · this result is deterministic: the same two personas on this problem always resolve the same way.