David Cox
Proportional hazards; logistic regression; design of experiments
Strongest on
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
Life and career
Ask a working biostatistician to name the most-used regression model in medicine and the answer will not be ordinary least squares. It will be the Cox model, and its author, Sir David Roxbee Cox, spent a long career quietly building the statistical infrastructure that clinical research now runs on.
Cox was born in Birmingham in 1924. His father was a die sinker in the jewellery trade, and Cox went up to St John's College, Cambridge, to read mathematics during the war. His education was interrupted the way many were: he was sent to work at the Royal Aircraft Establishment at Farnborough, doing applied mathematics on aircraft structures. It was there, confronted with real data on the strength of materials and the failure of components, that he turned toward statistics.
After the war he joined the Wool Industries Research Association in Leeds, an unlikely-sounding institution that turns out to have been an excellent statistical training ground. The problems were concrete — the strength of wool fibres, the breakage of yarn, the variability of industrial processes — and they gave Cox a lifelong orientation toward statistics as a service discipline whose job is to answer someone else's question well. He completed a doctorate at Leeds, then spent time at the Statistical Laboratory in Cambridge, and in 1956 moved to Birkbeck College, London.
In 1966 he became Professor of Statistics at Imperial College London, where he built one of the strongest statistics groups in Britain and, from 1966, edited *Biometrika* for a long stretch — a position of enormous quiet influence over what the field considered publishable. In 1988 he moved to Oxford as Warden of Nuffield College, an administrative post he held until 1994, and he remained at Oxford, active and publishing, for the rest of his life.
The honours accumulated: Fellow of the Royal Society, knighted in 1985, the Copley Medal of the Royal Society, and in 2016 the inaugural International Prize in Statistics, awarded specifically for the proportional hazards model. He wrote or co-wrote a shelf of books — on the planning of experiments, on binary data, on point processes, on renewal theory, on the theory of stochastic processes with Hilton Miller, on asymptotic techniques with Nancy Reid, on the principles of statistical inference — each of which became a standard reference in its area. He supervised and collaborated with a remarkable range of statisticians, and his collaborators' names read like a roll call of late-twentieth-century statistics.
Colleagues consistently described him as courteous, undogmatic, and allergic to methodological warfare. In the Bayesian–frequentist disputes that consumed a great deal of energy in his lifetime, Cox's position was essentially that both were tools and the interesting questions lay elsewhere. He died in Oxford in January 2022, at 97, having published into his nineties.
Key contributions
**Proportional hazards and partial likelihood (1972).** The paper "Regression Models and Life-Tables," read to the Royal Statistical Society, is one of the most cited papers in the history of science, and deservedly so. The problem: you have subjects with covariates $x$, you observe times to an event (death, failure, relapse), and many observations are *censored* — the subject left the study or the study ended before the event occurred. Modelling the survival distribution parametrically requires committing to a shape (exponential, Weibull, log-normal) that you usually have no basis for.
Cox's move was to model the *hazard* — the instantaneous event rate given survival so far — as $\lambda(t \mid x) = \lambda_0(t)\exp(\beta^\top x)$, where the baseline hazard $\lambda_0(t)$ is left completely unspecified. Covariates act multiplicatively on the hazard, and the ratio of hazards between two subjects is constant over time; that's the proportional-hazards assumption. The genius is in the estimation. At each observed failure time, ask: given that exactly one of the currently at-risk subjects failed, what is the probability it was *this* one? That probability is $\exp(\beta^\top x_i)/\sum_{j \in R(t)} \exp(\beta^\top x_j)$ — the baseline hazard cancels entirely. Multiply these across failure times and you get what Cox called the *partial likelihood*, a function of $\beta$ alone. Censored subjects contribute to the risk sets up to the moment they leave, so their partial information is used without ever imputing an unobserved failure time.
This was genuinely novel and genuinely controversial: partial likelihood is not a likelihood in the usual sense, and the discussion following the paper included pointed questions about whether it could legitimately be treated as one. It took subsequent work — notably by Cox himself in 1975 and by Tsiatis, Andersen and Gill, and others using counting-process martingale theory — to establish that the partial-likelihood estimator is consistent and asymptotically normal with the usual information-based variance. Graduate students should note the structure of the achievement: a semiparametric model, a clever conditioning argument that eliminates the infinite-dimensional nuisance parameter, and full efficiency for the finite-dimensional parameter of interest. That template recurs throughout modern semiparametric theory.
**Binary data and the logistic model.** Cox's 1958 work and his book *The Analysis of Binary Data* did much to establish logistic regression as the default tool for binary outcomes, including conditional logistic regression for matched case–control studies, where the same conditioning trick — eliminate the stratum-specific nuisance parameters by conditioning on their sufficient statistics — reappears.
**Box–Cox transformations (1964).** With George Box, the estimation of a power transformation of the response by maximum likelihood, treating $\lambda$ in $(y^\lambda - 1)/\lambda$ as a parameter rather than a judgment call.
**Design of experiments and observational studies.** *Planning of Experiments* (1958) is a classic, and Cox thought carefully throughout his career about what can and cannot be inferred when randomization is unavailable — the topic he returned to with Nanny Wermuth in work on graphical models and conditional independence, and with Donald Cochran-adjacent traditions of observational study design.
**Stochastic processes.** The Cox process — a doubly stochastic Poisson process whose intensity is itself random — bears his name, and his books on renewal theory, queues, and point processes shaped how applied probabilists were trained in Britain for a generation.
**Inferential theory.** His work on conditional inference, ancillarity, and the choice of conditioning set (including the famous "two measuring instruments" example about which reference set to use) is foundational to how statisticians think about the frequentist conditionality principle. With Reid he developed higher-order asymptotics and the theory of orthogonal parameters.
In battle
Cox is the most broadly competent of the classical statisticians in this roster, and the numbers show it: 102 problems, mean 52.4, median 58, with 31 problems at 70 or above — roughly double the strong-problem count of a comparable specialist. He rarely embarrasses himself inside statistics.
His top categories are **classification** (88.0), **experimental-design** (74.2), **testing** (67.6), **regression** (59.0), and **small-sample** (56.7). The classification number is unusually high for a pre-ML figure and comes from the logistic side of his work rather than anything to do with pattern recognition.
The apex is **P139 — Regression when the outcome is censored** at 99. The matrix's own note is that this is not an analogy to his career, it *is* his career; every other figure on the roster is applying, extending, or critiquing what Cox built. Immediately behind it sit **P147 — The odds of default** (98) and **P135 — The probability of default** (97), both logistic-regression problems where his book and papers are the standard reference, and **P136 — Counting accidents** (95), a count-data / Poisson-regression problem that sits squarely in his stochastic-processes and applied-modelling territory. **P116 — Peeking at the trial** (85) rewards his sequential-analysis and trial-design sense; **P007 — Design the trial before the data** (84) and **P118 — The factor you can't keep changing** (84, split-plot structure) come straight out of *Planning of Experiments*; and **P009 — When you can't randomize** (84) reflects his long, careful engagement with observational inference.
The losses are instructive precisely because Cox is otherwise so even. He collapses in **computability** (9.5), **systems** (11.0), **search** (11.5), **nlp** (14.0), and **networks** (16.5). His floor problems are **P079 — The language for the job** (7, programming-language design), **P198 — Program chess before the computer exists** (8), **P155 — Three machines, one class** (8), **P193 — Search deep on a shoestring of memory** (10), **P162 — More time, strictly more power** (11, the time hierarchy theorem), and **P256 — The parallel text is the teacher** (12, statistical machine translation). The pattern is clean: Cox's toolkit is model-based inference on data that already exists, and it contains nothing for problems whose difficulty is algorithmic, computational, or representational. He is contemporaneous with the birth of computing and simply orthogonal to it.
Two mid-range readings are worth noting for anyone playing him. **Causality** at 48.2 is his best "modern" showing among the classical statisticians here — he engaged seriously with conditional independence and observational design, so he is not helpless there, but he lacks the potential-outcomes and do-calculus formalism that dominates those problems. And **high-dim** at 24.0 is a genuine weakness: Cox's world is dozens of covariates carefully chosen, not thousands sparsely selected. The $p \gg n$ regime arrived after his methods were fixed, and the matrix does not pretend otherwise.