AI History Battle
Peeking at the trial experimental-design

It is the 1960s, and a clinical trial faces an ethical knife-edge: if the treatment is clearly working — or clearly harming — it is wrong to keep enrolling patients to the end, yet every interim peek at the accumulating data inflates the chance of a false positive, because repeatedly testing the same growing dataset is multiple testing in disguise. Design a monitoring plan with pre-specified interim analyses that can stop early for benefit or harm while keeping the overall error rate at its nominal level, spending your significance budget across the looks. Get it wrong and either patients are harmed by a trial that ran too long, or a fluke at an early peek halts a trial and crowns a treatment that does nothing.

sequentialmultiplicityerror control
b. 1965
tapped
55

Chose Bayesian computation and model checking — right call.

Gelman's applied Bayesian workflow, developed from the 1990s onward and embodied in Stan, offers a genuinely different but coherent answer to interim monitoring: rather than spending a fixed frequentist error budget across looks, a Bayesian posterior can be updated continuously as data arrive and decisions made by posterior probability thresholds, sidestepping the multiple-testing correction the problem describes as a frequentist headache. His hierarchical models also naturally pool information across interim looks. The deductions are substantial: his approach is philosophically distinct from the alpha-spending framework the 1960s clinical-trial establishment actually adopted and regulators still largely require, his major work arrives three decades after the scene, and he has written more about general Bayesian workflow and model checking than about the specific regulatory history of sequential clinical trial design.

was tapped · ask the professor
0

The professor reaches for his one statistics lecture slide on p-values, forgets which direction Type I error goes, and by the time he has found the chalk, Wald has already run the sequential probability ratio test by hand, stopped the trial early, and gone to lunch. He teaches AI at Berkeley and consults for companies that pay him to know things like this, which is precisely why this scoring system exists: to remind him, publicly, in front of his own students, that reading about the history of statistics is not the same as being able to do it under 1960s wartime pressure with a slide rule. He loses to Wald. He loses to Cox. He would probably lose to the placebo arm.

Head to head 20 over 2 battles
Read Gelman Read Santerre Leaderboard

Battle #20 · 8/9/2026, 5:07:45 PM · this result is deterministic: the same two personas on this problem always resolve the same way.