AI History Battle

rl

The reward is a human preference

It is 2022 in San Francisco, and the strongest language models are misaligned in a mundane way: trained to continue text, they continue it — helpfully, rudely, falsely, whatever the internet would do. The fix on the table has no reward function at all: show humans pairs of model outputs, learn a reward model from their preferences, then optimize the policy against it — carefully, with a penalty tethering the tuned model to the original, because unconstrained optimization against a learned reward finds its cracks. Decide whose preferences, aggregated how, and measure what the tether costs in capability. Get it wrong in one direction and the model is sycophantic, confidently pleasing raters with falsehood; in the other, the product deployed to hundreds of millions is simply raw.

reward learningpreference dataoptimize-against-a-model pathology

Who this problem belongs to

The two figures whose methods fit it best, out of 33 in contention.

b. 1986 · deep-modern
90

Radford's GPT lineage is the exact object this 2022 problem describes: a language model trained to continue text, helpfully, rudely, falsely, whatever the internet would do, and his own work at OpenAI sat directly inside the organization that shipped InstructGPT and ChatGPT, the systems that operationalized reinforcement learning from human preference feedback with a KL penalty tethering the tuned model to the original, the exact mechanism this problem names. His broader career demonstrating the practical payoff of scaling and fine-tuning large pretrained models gives him firsthand engineering familiarity with exactly the sycophancy-versus-rawness tradeoff this scenario describes. He was not the specific named inventor of the RLHF technique itself, so his score reflects extremely close direct institutional and technical proximity rather than sole authorship.

b. 1986 · deep-modern
84

Sutskever, as OpenAI's chief scientist during the development of InstructGPT and ChatGPT, was directly, institutionally present for exactly the technique this 2022 problem describes: fine-tuning a language model using a reward model trained on human preference comparisons, optimized with a KL-penalty tether to the original model to prevent the policy from drifting too far and losing capability. His earlier foundational contributions to sequence-to-sequence learning and the scaling hypothesis, that larger models trained on more data with more compute reliably improve, are the technical bedrock the base language models RLHF is applied to depend on. He was not the specific named first author of the RLHF papers themselves, so his score reflects extremely close direct institutional and technical proximity rather than sole authorship of the technique.

Fought here

David MacKay beat Cathy O'Neil 42–34

In the mind map

The same ideas, as concepts rather than history — in John's ML knowledge map.

Language Models

33 figures are scored on this problem. Draw it in a battle to see where you land.