AI History Battle
Peeking at the trial experimental-design

It is the 1960s, and a clinical trial faces an ethical knife-edge: if the treatment is clearly working — or clearly harming — it is wrong to keep enrolling patients to the end, yet every interim peek at the accumulating data inflates the chance of a false positive, because repeatedly testing the same growing dataset is multiple testing in disguise. Design a monitoring plan with pre-specified interim analyses that can stop early for benefit or harm while keeping the overall error rate at its nominal level, spending your significance budget across the looks. Get it wrong and either patients are harmed by a trial that ran too long, or a fluke at an early peek halts a trial and crowns a treatment that does nothing.

sequentialmultiplicityerror control
b. 1991
tapped · ask the professor
15

Abebe's work on algorithms and inequality, and mechanism design for social good, engages deeply with how quantitative systems affect vulnerable populations -- a values-alignment relevant to this problem's ethical stakes around patient welfare in a trial. But her technical contributions concern algorithmic fairness and resource-allocation mechanism design in social contexts (housing, homelessness services), not sequential statistical testing, error-rate control, or clinical trial monitoring methodology. There is no direct methodological bridge from her body of work, which is contemporary (2010s-2020s) and computer-science-adjacent, to the specific alpha-spending and stopping-boundary apparatus this 1960s clinical-trial problem requires; her relevance here is limited to a general ethical sensibility about the stakes of getting quantitative decisions about people wrong.

was tapped · ask the professor
0

The professor reaches for his one statistics lecture slide on p-values, forgets which direction Type I error goes, and by the time he has found the chalk, Wald has already run the sequential probability ratio test by hand, stopped the trial early, and gone to lunch. He teaches AI at Berkeley and consults for companies that pay him to know things like this, which is precisely why this scoring system exists: to remind him, publicly, in front of his own students, that reading about the history of statistics is not the same as being able to do it under 1960s wartime pressure with a slide rule. He loses to Wald. He loses to Cox. He would probably lose to the placebo arm.

Head to head 10 over 1 battle
Read Abebe Read Santerre Leaderboard

Battle #46 · 8/9/2026, 8:36:54 PM · this result is deterministic: the same two personas on this problem always resolve the same way.