experimental-design
Peeking at the trial
It is the 1960s, and a clinical trial faces an ethical knife-edge: if the treatment is clearly working — or clearly harming — it is wrong to keep enrolling patients to the end, yet every interim peek at the accumulating data inflates the chance of a false positive, because repeatedly testing the same growing dataset is multiple testing in disguise. Design a monitoring plan with pre-specified interim analyses that can stop early for benefit or harm while keeping the overall error rate at its nominal level, spending your significance budget across the looks. Get it wrong and either patients are harmed by a trial that ran too long, or a fluke at an early peek halts a trial and crowns a treatment that does nothing.
Who this problem belongs to
The two figures whose methods fit it best, out of 37 in contention.
This problem is Wald's own invention transplanted a decade forward. His sequential probability ratio test (1943-45), developed for the U.S. Navy to shorten wartime munitions inspection, is precisely a monitoring rule that accumulates evidence and stops as soon as it crosses a pre-specified boundary for accepting or rejecting a hypothesis, while provably controlling both error rates at their nominal levels no matter when the trial stops. That is the exact engineering task the problem poses: pre-specified boundaries, controlled overall error, and the ethical benefit of stopping early. His decision-theoretic framing of loss and risk also supplies the language for weighing patient harm against false conclusions. The only reason he is not a perfect 100 is that clinical group-sequential boundaries (alpha-spending, interim looks at fixed calendar points rather than continuous monitoring) were refined by his successors after his 1950 death.
Cox spent his career as one of the era's leading authorities on clinical trial methodology -- his proportional hazards model (1972) became the standard tool for survival analysis in exactly this kind of trial, and his broader work on sequential and adaptive methods in medical statistics, alongside decades advising British trial-design practice, put him at the center of the group-sequential-methods conversation as it matured through the 1970s-80s. He would recognize immediately that repeated interim looks are 'multiple testing in disguise' and know the corrective machinery (spending functions, boundary adjustments) that followed Wald's original insight. He sits just below Wald and Neyman because he is the era's premier synthesizer and practitioner of trial methodology rather than the originator of the sequential-testing mathematics itself, which predates his most active period by two decades.
Fought here
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
37 figures are scored on this problem. Draw it in a battle to see where you land.