Karl Pearson
Correlation, chi-squared; founder of mathematical statistics
Played by Karan
Strongest on
Battles
The robot in the warehouse W Risi Kondor
Unroll the swiss roll W Raj Reddy
Choose the first hundred believers L Josh Tenenbaum
Which examples deserve labels? W Josh Tenenbaum
Why least squares, exactly?
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
Life and career
Karl Pearson was born Carl Pearson in London in 1857, into a Quaker family of some means, and changed the spelling of his own first name in his twenties — an act of homage to Karl Marx, whose work he had encountered while studying in Germany. That detail is a useful entry point, because Pearson's intellectual life was never confined to statistics and was never politically neutral.
He read mathematics at King's College, Cambridge, graduating as Third Wrangler in 1879, then spent time at Heidelberg and Berlin studying physics, metaphysics, German literature, and law. He was called to the bar and did not practice. He wrote on medieval German folklore, on the history of religion, on Spinoza, and on the position of women — he lectured to radical clubs, helped found a Men's and Women's Club for the discussion of relations between the sexes, and held socialist and freethinking commitments throughout. In 1884 he took the chair of applied mathematics and mechanics at University College London, and in 1892 published *The Grammar of Science*, a widely read positivist manifesto arguing that science is a method rather than a body of content, that all its claims reduce to descriptions of sense-impressions, and that no field is scientific except insofar as it applies that method. Einstein's reading group is said to have studied it.
The turn to statistics came from two colleagues. Francis Galton had published *Natural Inheritance* in 1889 with its observation of regression toward the mean; and Walter Frank Raphael Weldon, a zoologist at UCL, brought Pearson biological measurement data — crab dimensions, shrimp populations — with questions Pearson had the mathematics to answer. Between roughly 1893 and 1912 Pearson produced the series *Mathematical Contributions to the Theory of Evolution*, running to eighteen major papers, which essentially invented mathematical statistics as an autonomous discipline. He founded *Biometrika* with Weldon and Galton in 1901 as a journal for this work after the Royal Society proved unwilling to publish papers mixing biology and mathematics. Galton's bequest endowed a chair in eugenics at UCL, which Pearson took in 1911, folding his Biometric Laboratory into what became the first university statistics department in the world.
He was a formidable and difficult man. His feud with Ronald Fisher, beginning around 1917–20 and lasting until Pearson's death, was bitter, personal, and consequential — Pearson controlled *Biometrika* and used it, Fisher retaliated in kind, and the field's development was shaped by their inability to be in a room together. The eugenics is not a footnote and should not be treated as one. Pearson's *Annals of Eugenics* and the laboratory he ran produced work explicitly aimed at questions of national heredity, and his 1925 study with Margaret Moul on Jewish immigrant children in London argued against admitting them on hereditary grounds. This was mainstream in his milieu and it was also wrong on the evidence and monstrous in its use, and the statistical apparatus he built was in part built to serve it. He retired in 1933; UCL split his position into a statistics chair (to Fisher, of all people) and a eugenics chair (to his son Egon). He died in 1936. In 2020, UCL removed his name from buildings and lecture theatres.
Key contributions
**The correlation coefficient.** Pearson's r — the covariance of two variables divided by the product of their standard deviations — was given its modern definition and its distribution theory in his 1896 paper, building on Galton's more intuitive work. This is the single most-used number in applied quantitative science, and Pearson's contribution was to make it a properly estimated parameter with sampling behavior rather than a graphical impression. He extended it to partial and multiple correlation, giving the apparatus for asking about association between two variables holding others fixed, and to biserial and tetrachoric coefficients for the case where the underlying continuous variables are observed only in categories.
**Regression toward the mean, explained.** Galton observed that tall fathers have sons closer to the population mean and reached for a biological force to explain it. Pearson supplied the mathematics showing that the regression slope is a function of the correlation, and that regression toward the mean is a necessary consequence of imperfect correlation between any two variables — not a force at all. Every discussion of regression to the mean in program evaluation, sports analytics, and medical trials descends from this clarification, and getting it wrong remains one of the most common errors in applied work.
**The chi-squared goodness-of-fit test (1900).** This is arguably his greatest single paper. Given observed counts in categories and expected counts under a hypothesized distribution, the statistic Σ(O−E)²/E has, in large samples, a chi-squared distribution whose degrees of freedom depend on the number of categories. For the first time there was a general, distribution-free procedure for asking whether a model fits data *at all*, rather than just estimating parameters within an assumed model. It is the ancestor of every goodness-of-fit test, of the chi-squared test of independence in contingency tables, and of the asymptotic chi-squared reference distribution used by likelihood ratio, Wald, and score tests throughout modern statistics. Fisher's famous correction to the degrees of freedom when parameters are estimated from the data was one of the sparks of their long war.
**The method of moments and the Pearson system of curves.** Faced with biological data that was visibly skewed and non-normal, Pearson defined a family of distributions generated as solutions to a differential equation, parameterized so that its members cover a wide range of skewness and kurtosis — the family includes the normal, beta, gamma, Student's t, and others as special cases. He fitted them by matching sample moments to theoretical moments. Method of moments is less efficient than maximum likelihood and Fisher would demolish it on those grounds, but it is computationally trivial, requires no iteration, and remains the standard route to initial estimates and to identification arguments in latent-variable and tensor-decomposition methods today.
**Principal component analysis (1901).** Pearson's paper "On Lines and Planes of Closest Fit to Systems of Points in Space" poses the problem of finding the lower-dimensional affine subspace minimizing perpendicular distance to a cloud of points, and gives the solution in terms of the principal axes of the scatter. Hotelling later developed the variance-maximization formulation and the name, but the geometric origin is Pearson's, and PCA remains the most widely used dimension-reduction method in existence.
**Institutions.** *Biometrika*, the Biometric Laboratory, the first statistics department, the training of a generation, the standard vocabulary — standard deviation, histogram, and much else are his coinages. Pearson made statistics a profession.
In battle
Pearson's numbers are those of a broad, reliable heavyweight rather than a narrow specialist: mean 36.5, median 34, five problems above 80, nine above 70, and only twenty-six at or below 20 — a notably low weakness count for a figure of his era. He shows up competently in more places than Bayes, Gauss, or Markov do.
His home turf is measurement of association and assessment of fit. "Why tall fathers have shorter sons" (99) is his own discovery with Galton, on that exact dataset — he is not applying an outside toolkit, he built the toolkit for the problem. "Does the model fit at all?" (97) is the 1900 chi-squared paper. "Two proportions, tiny cells" (86) is the contingency-table problem where his chi-squared apparatus meets small counts, and "The ruler that lies a little" (84) is measurement error and attenuation of correlation — the reliability theory that follows directly from his coefficient. "The recombination fraction from a small cross" (80) and "Counting yeast in the pitching square" (72) are biological estimation problems of exactly the kind his moment methods and curve system were designed for. "Thirty percent chance of rain" (78) is calibration and goodness-of-fit for probabilistic forecasts. "Ten patients, one rare disease" (72) is small-sample proportion estimation. His category profile is correspondingly wide: high-dimensional 54.0 (PCA), small-sample 53.7 across fifteen problems, testing 49.9 across fourteen, regression 43.6 across sixteen, classification 40.0.
Two things are worth flagging about the middle of his range. His causality average is 33.1 across sixteen problems — mediocre, and by design. Pearson was philosophically committed to the position that correlation is all there is, that causation is a metaphysical residue science should discard, and he said so at length in *The Grammar of Science*. That commitment is precisely what caps him on causal problems: he has the best association machinery of his generation and a principled refusal to take the next step. His fairness average of 32.8 is unexpectedly respectable, and a student should sit with the discomfort of that: the man built the measurement apparatus that modern fairness auditing runs on, and also put it to eugenic use.
His losses are uniformly computational or logical. "Agreement among the unreliable" (8) is inter-rater agreement and consensus among noisy labelers — Cohen's kappa territory, later and with different foundations. "Prove the program correct" (7) and "Three machines, one class" (7) are program verification and machine-model equivalence, pure computability with a systems average of 6.5 and computability average of 6.5 to match. "Is there a fast route through every city?" (6) is NP-hardness and combinatorial complexity. "A computer shared by fifty" (5) is time-sharing operating systems, and "Optimize across the datacenter" (4) is distributed consensus optimization — his floor, where the profile notes his entire method assumes a single centralized dataset.
Play Pearson on association, goodness-of-fit, contingency tables, dimension reduction, measurement reliability, and any problem where the question is *how well does this model describe this data*. He will hold his own almost anywhere in classical statistics and lose completely the moment the problem becomes about machines, algorithms, or proof.