It is 2023, and a lawyer has just been sanctioned for filing a brief containing six precedents that do not exist — invented by a language model, complete with plausible citations, in prose indistinguishable from the real thing. The model was not malfunctioning; it was doing exactly what next-token prediction trains it to do. Diagnose the failure honestly: why maximum-likelihood training on text rewards plausibility rather than truth, why the errors concentrate exactly where users cannot check, and what combination of retrieval grounding, calibration, and honest abstention actually reduces fabrication rather than hiding it. Then design the evaluation that measures it. Get it wrong and fluent fabrication ships into medicine, law, and journalism at scale — trusted precisely because it reads like expertise.
Nowak's work on active learning and sparse recovery, developed mainly through the 2000s, addresses how to efficiently select informative examples and extract signal from high-dimensional data, general statistical learning concerns with only a loose, indirect connection to why a large language model confidently fabricates legal citations rather than expressing appropriate uncertainty. His research on statistical signal processing meeting machine learning has no direct bearing on generative text factuality, calibration for language models, or legal-domain deployment specifically. Nowak's actual technical contributions target statistical estimation and query-selection strategies for machine learning systems developed in domains far removed from large-scale generative language modeling. His relevance to this problem is essentially peripheral, connected only through the broadest shared category of statistical machine learning methodology.
Pearl's causal revolution, culminating in his do-calculus and the 2018 book The Book of Why, gives this problem its sharpest available diagnostic tool: a rigorous account of why models trained to predict correlational patterns in text — what word plausibly follows another — have no mechanism for distinguishing a causally grounded true claim from a statistically fluent fabrication, since likelihood-based training never represents the difference between correlation and truth in the first place. His decades of work on Bayesian networks, developed from the 1980s onward, gave AI its first rigorous framework for representing uncertainty and causal structure explicitly, precisely what next-token prediction discards. Pearl worked in causal inference and probabilistic reasoning generally, not language models or legal-citation fabrication specifically, and did not address the transformer architecture directly.
Battle #132 · 8/10/2026, 11:39:21 AM · this result is deterministic: the same two personas on this problem always resolve the same way.