testing
Weight of evidence at Bletchley
It is 1941 at Bletchley Park, and each intercepted German message is a scrap of evidence for or against a guessed rotor setting of the Enigma — but a single scrap decides nothing, and testing every setting exhaustively is hopeless. Devise a way to accumulate evidence: assign each observation a weight in favor of a hypothesis, add these weights as messages arrive, and stop when the total crosses a threshold that guarantees your error rates — deciding for or against a setting with as little intercepted material as possible. Get it wrong and you either commit to a wrong setting and misread a day's traffic, or demand so much evidence that the intelligence is stale before you act — and in 1941 the latency is measured in sunk ships.
Who this problem belongs to
The two figures whose methods fit it best, out of 55 in contention.
This is, almost literally, Turing's actual war. At Bletchley Park from 1939 he led Hut 8's attack on naval Enigma, and by 1941 he had built Banburismus, a sequential procedure that assigned each intercept a numerical 'weight of evidence' in decibans, summed those weights as messages arrived, and stopped as soon as the accumulated total crossed a threshold implying a settled rotor guess. It let cryptanalysts commit to or discard a wheel order with far less material than exhaustive testing demanded, directly shortening the delay before a day's U-boat traffic could be read. He derived the deciban scale himself, working with Jack Good. No other figure on this list did this exact job, in this exact building, in this exact year.
Wald never worked at Bletchley, but he solved the identical mathematical problem from the American side at almost the same moment. As a member of the Statistical Research Group from 1943, he formalized the sequential probability ratio test: accumulate a log-likelihood ratio observation by observation, stop when it crosses one of two thresholds set by the error rates you are willing to tolerate, and prove that this procedure is optimal in expected sample size among all tests achieving those error rates. That theorem is the rigorous backbone underneath Turing's more improvised deciban bookkeeping. He scores just below Turing only because his work, however mathematically superior, was not this room, this war, or these rotors — it was the general theory Bletchley's practice anticipated.
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
55 figures are scored on this problem. Draw it in a battle to see where you land.