AI History Battle

It is 2018, and one Atari game has become the field's public humiliation: Montezuma's Revenge, where the first reward sits beyond ladders, ropes, a key, and a locked door — hundreds of correct actions with zero feedback — and the agents that conquered every other game score nothing. Random exploration will not stumble through; the agent needs reasons to act before the world supplies any. Build intrinsic motivation that is not a bug generator: curiosity about the unpredicted, novelty bonuses over learned state abstractions — and confront the noisy-TV trap, where pure randomness is infinitely 'novel' and the curious agent stares at static forever. Get it wrong and RL remains a method for dense-reward games only, useless wherever feedback is rare and delayed — which is everywhere that matters.

sparse rewardintrinsic motivationexploration pathologies
b. 1970
tapped
18

Candes's compressed sensing work, recovering signals from surprisingly few measurements given the right sparsity assumptions, has a thin conceptual echo in this problem's need to extract useful exploratory signal from an environment that offers almost no reward feedback, both problems are about recovering structure from scarce information. But his actual published research is mathematical signal processing and high-dimensional statistics, not reinforcement learning, exploration strategy, or intrinsic motivation, and nothing in his career touches sparse-reward sequential decision-making or curiosity-driven agents specifically, making this a distant structural analogy rather than a demonstrated methodological connection to how Montezuma's Revenge was actually eventually solved by DeepMind and Berkeley researchers working in reinforcement learning specifically.

b. 1974
was tapped · ask the professor
15

Vidal's work on subspace clustering and generalized PCA addresses extracting low-dimensional structure from high-dimensional visual data, a loosely relevant capability for building the learned state abstractions this problem's novelty-bonus systems rely on to compress raw Atari pixels into something more tractable. But his research career is centered on computer vision and geometric data analysis, not reinforcement learning, reward structure, or exploration strategy, and nothing in his published work engages sparse-feedback sequential decision problems, curiosity bonuses, or the noisy-TV trap specifically, making this a distant structural analogy from a body of work built for an entirely different application than sequential, reward-driven decision-making under sparse and severely delayed feedback signals over long time horizons.

Head to head 20 over 2 battles
Read Candes Read Vidal Leaderboard

Battle #89 · 8/10/2026, 11:36:53 AM · this result is deterministic: the same two personas on this problem always resolve the same way.