AI History Battle

It is 2022 in San Francisco, and the strongest language models are misaligned in a mundane way: trained to continue text, they continue it — helpfully, rudely, falsely, whatever the internet would do. The fix on the table has no reward function at all: show humans pairs of model outputs, learn a reward model from their preferences, then optimize the policy against it — carefully, with a penalty tethering the tuned model to the original, because unconstrained optimization against a learned reward finds its cracks. Decide whose preferences, aggregated how, and measure what the tether costs in capability. Get it wrong in one direction and the model is sycophantic, confidently pleasing raters with falsehood; in the other, the product deployed to hundreds of millions is simply raw.

reward learningpreference dataoptimize-against-a-model pathology
b. 1972
tapped · ask the professor
42

Chose Feedback-loop spotting — wrong. Outcome auditing over input inspection was the one that fit.

O'Neil's Weapons of Math Destruction is precisely an argument that systems optimized against a proxy metric, exactly what a learned reward model is, will exploit the gap between the proxy and the real goal unless independently and continuously audited, the exact pathology this problem's "unconstrained optimization against a learned reward finds its cracks" warning describes. Her broader accountability framework directly engages this problem's demand to decide whose preferences are aggregated and how, a question with real distributional consequences her work has repeatedly highlighted in other algorithmic domains. She did not build the RLHF technique or work on large language models specifically, so her relevance is a directly applicable accountability lens rather than technical authorship of the reward-modeling mechanism itself.

1967–2016
was tapped
34

MacKay's information-theoretic treatment of inference and his Bayesian neural network research treated calibrated, honest uncertainty as central to trustworthy machine learning, directly relevant to this problem's demand to measure what the KL-penalty tether costs in capability, a question about honestly characterizing a tradeoff rather than hiding it. His information-theoretic vocabulary is also the natural language for describing the KL-divergence penalty this problem's RLHF mechanism explicitly uses. He worked mainly in information theory and Bayesian machine learning inference rather than reinforcement learning from human feedback or large language models specifically, and died in 2016 before RLHF was formalized, so his relevance is a rigorous theoretical vocabulary rather than direct historical engagement with this 2022 technique.

Head to head 02 over 2 battles
Read O'Neil Read MacKay Leaderboard

Battle #40 · 8/9/2026, 7:13:53 PM · this result is deterministic: the same two personas on this problem always resolve the same way.