AI History Battle
Engraved card portrait of Kiyosi Ito

Kiyosi Ito

1915–2008 · early-stat

Stochastic calculus; the Ito integral

0wins
0losses
win rate

Strongest on

99 Calculus for a jagged path 97 The fortune that never sits still 67 Bet with information theory 49 Roll the dice at Los Alamos 45 Dynamic programming's curse 45 Catch the process the moment it drifts

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Langevin Dynamics Bellman Equation Diffusion Models Brownian Motion

Life and career

Kiyosi Itô built the mathematical machinery that runs modern finance, filtering, and diffusion modeling, and he did the foundational work in wartime Japan, cut off from the international mathematical community, while employed at a government statistics bureau.

He was born in 1915 in Hokusei-cho, in Mie Prefecture. He entered the University of Tokyo in 1935 to study mathematics, at a moment when probability theory was regarded in Japan — as in much of the world — as a somewhat disreputable applied subject rather than real mathematics. Kolmogorov's 1933 axiomatization had just appeared and was beginning to change that, and Itô read it. What captured him was the work of Paul Lévy on the paths of stochastic processes, and of Kolmogorov and Feller on the analytic side. Lévy's writing was notoriously intuitive and difficult to follow rigorously, and Itô later said, in effect, that his own project was to make Lévy's insights into theorems.

He graduated in 1938 and took a position at the Cabinet Statistics Bureau, where he worked for the next several years. This was not a research post; it was government statistical work, and he did mathematics around it. His two foundational papers on stochastic integration appeared in 1942 and 1944 — the 1942 one, on what he called "differential equations determining Markoff processes," was published in Japanese in a journal that essentially nobody outside Japan could read during the war, and its content did not reach the West until afterward. He completed his doctorate at Tokyo in 1945.

After the war he moved to Nagoya University, then in 1952 to Kyoto University, where he remained for the rest of his career. He spent extended periods abroad — at the Institute for Advanced Study in Princeton in the mid-1950s, at Stanford, at Aarhus, at Cornell — as his work became known and then became central. He was director of the Research Institute for Mathematical Sciences at Kyoto. He collaborated with Henry McKean on *Diffusion Processes and Their Sample Paths* (1965), the book that consolidated the theory for a generation, and developed excursion theory, a way of decomposing a Markov process's path into its excursions away from a point, which is now standard.

Colleagues consistently described him as courteous, unassuming, and slightly bemused by the enormous applied industry that grew out of his work. He received the Wolf Prize in 1987, the Kyoto Prize in 1998, the inaugural Gauss Prize in 2006 — an award created specifically for mathematics with outsized impact outside mathematics, and awarded to him for exactly that reason — and Japan's Order of Culture. His daughter accepted the Gauss Prize on his behalf; he was ninety and in frail health. He died in Kyoto in 2008.

Key contributions

**The Itô integral.** The problem is genuinely a scandal, and it is worth stating precisely. Brownian motion has sample paths that are continuous everywhere and differentiable nowhere; worse, they have unbounded total variation on every interval. This means the Riemann–Stieltjes integral ∫ f dB simply does not exist path by path — the standard construction requires the integrator to have bounded variation, and Brownian motion does not. Yet the objects physics and probability most needed to describe were exactly these: a quantity buffeted continuously by noise.

Itô's construction sidesteps path-by-path definition entirely and builds the integral as an L² limit. For simple predictable integrands — piecewise constant, with each piece's value determined by information available *at the start* of its interval — define the integral as the obvious sum. The **Itô isometry**, E[(∫ H dB)²] = E[∫ H² dt], shows this map is an isometry from a space of predictable processes into L², so it extends by continuity to a much larger class of integrands. The result is a well-defined stochastic integral that is itself a martingale.

The choice to evaluate the integrand at the *left* endpoint is not a technicality; it is the whole thing. It makes the integral non-anticipating — the integrand cannot peek at the increment it multiplies — which is what makes the martingale property hold and what makes the object meaningful for modeling a system that reacts to noise as it arrives. Stratonovich's alternative convention, evaluating at the midpoint, gives ordinary chain-rule calculus but loses the martingale property; the two are related by an explicit correction and each is appropriate for different purposes.

**Itô's lemma.** This is the chain rule for stochastic calculus and the single most-used formula in the subject. If X satisfies dX = μ dt + σ dB and f is twice continuously differentiable, then

df(X) = f'(X) dX + ½ f''(X) σ² dt.

The second-order term is the surprise, and its origin is exact: Brownian motion accumulates quadratic variation at rate one, so (dB)² = dt rather than vanishing as it would in ordinary calculus. Second-order terms that classical Taylor expansion discards as negligible are, here, first order. Everything distinctive about stochastic calculus is that correction term. It is why the exponential of a Brownian motion has a drift correction, why the Black–Scholes PDE has the ½σ²S²∂²V/∂S² term, and why a naively transformed SDE gives the wrong answer.

**Stochastic differential equations.** Itô proved existence and uniqueness of strong solutions to dX = b(X,t) dt + σ(X,t) dB under Lipschitz and growth conditions, giving a rigorous pathwise construction of diffusion processes. This connected the analytic view (Kolmogorov's forward and backward equations for transition densities) to the pathwise view (a trajectory driven by noise), and the Feynman–Kac formula ties both to PDEs.

**Where the field uses it.** Mathematical finance in its entirety — Black–Scholes, term-structure models, risk-neutral pricing via Girsanov's change of measure. Nonlinear filtering: the Kalman–Bucy filter and the Zakai equation. Stochastic control and the Hamilton–Jacobi–Bellman equation. Population genetics diffusion models. And, most relevantly for a current data-science audience, **score-based generative models and diffusion models**: the forward noising process is an SDE, the reverse-time generative process is a different SDE obtained by Anderson's time-reversal result, and Langevin dynamics sampling is an SDE discretization. When a diffusion model generates an image, it is integrating an Itô SDE.

In battle

Itô's profile is the most extreme specialization on the roster: mean 18.6, median 14, exactly one problem above 70, and seventy-one at or below 20. He has one perfect weapon and almost nothing else.

That weapon is "Calculus for a jagged path" at 99, and the profile is blunt that it is not an analogy — it is the actual historical event. Brownian motion is nowhere differentiable, ordinary calculus fails, and Itô built the integral and the correction term that fix it. This is the highest score he could earn and he earns it outright.

Below that, the drop is immediate and the pattern of his second tier is instructive. "Bet with information theory" (67) is the Kelly criterion — log-optimal betting, where the geometric-mean growth rate and the variance drag are exactly the Itô correction in a different costume, and any continuous-time treatment of the problem is his mathematics. "Roll the dice at Los Alamos" (49) is Monte Carlo simulation, adjacent to his stochastic-process machinery without being his method. "Dynamic programming's curse" (45) reaches continuous-time control, where the HJB equation sits directly downstream of his SDEs. "Contagion on the network" (45) is epidemic modeling, where diffusion approximations to branching and SIR processes apply. "Catch the process the moment it drifts" (45), "The adaptive dose-finder" (44), and "Stopping the sequential test" (44) are sequential and change-point problems — Wald's territory, where Itô scores respectably because the modern continuous-time treatment of optimal stopping is built on his calculus, but where he loses decisively to the person who actually invented the procedures.

His losses are total and reveal the boundary sharply. Three failure modes recur. **Applied causal inference and epidemiology**: "The paradox in the admissions data" (6) is Simpson's paradox, "The therapy the trial reversed" (6) is confounding by indication in the WHI trial. These are about study design and confounding structure, and no amount of stochastic analysis touches them. **Algorithms and discrete computation**: "The cluster that iterates" (5) is k-means, "A million parsed sentences" (4) is treebank-based statistical parsing. **Modern sociotechnical problems**: "The agent that games its reward" (4) is reward hacking and specification gaming, "The variable you removed is still there" (4) is proxy discrimination in fairness auditing — his floor, where the profile notes he would have essentially nothing specific to contribute. His classification average is 7.0, systems 6.5, computability 9.0, fairness 9.5.

The single-sentence summary of his battle identity is exact: continuous-time randomness is his, and discrete or combinatorial terrain is not. Every problem he wins involves a process evolving in continuous time under noise. Every problem he loses involves either a discrete structure, a designed study, an algorithm, or a human institution. A student should read his sheet as the clearest illustration on the roster of what a *deep but narrow* contribution looks like: one of the most consequential pieces of twentieth-century mathematics, worth 99 points on one problem and under 20 on seventy-one others.