№ 48 · machine learning

Learning from preferences

Ask someone to score an essay from one to ten and they hesitate. Show them two and ask which is better, and they answer at once. Learning from preferences builds a training signal from the easier question.

What it is

A model produces two outputs — two robot video clips, two answers to a question. A person picks one. Thousands of those choices fit a second model, a reward model, that scores any output. The first model is then trained to produce outputs it scores highly. Nobody writes down what "good" means; it is reconstructed from comparisons.

Why it matters

The wall was specification. Training a robot to backflip by reinforcement learning needs a reward function — a formula scoring each moment — and nobody can write one for a good backflip. Christiano and colleagues showed in 2017 that you can skip the formula: a person compares pairs of one-to-two-second clips, and the agent learns from feedback on less than 1% of its interactions. The Hopper robot learned a backflip from 900 comparisons, under an hour of a person's time. Their benchmark robots came close to ordinary reinforcement learning with 700 comparisons; Atari took 5,500.

Later the same recipe reached language. Ouyang and colleagues started from GPT-3, first fine-tuned it on answers written by about 40 contractors, then had those contractors compare the supervised models' answers, and trained against the resulting reward model. Human evaluators preferred the 1.3-billion-parameter tuned model to the original 175-billion-parameter GPT-3 — over 100 times its size. The 175B tuned version beat plain GPT-3 85 ± 3% of the time.

Interactive Six swatches have hidden darkness scores. Click the darker of each pair to act as the labeller, or press auto to add noisy comparisons — the Bradley–Terry fit re-runs on every click and the chart tracks how much of the true ranking it has recovered.

which is darker?
0 comparisons
A real fit, on illustrative data. Each swatch has a hidden score (its darkness, on an arbitrary scale); the six scores and the simulated labeller are made up for this demonstration, not taken from either paper. Auto comparisons are drawn from the Bradley–Terry model itself: the labeller picks A over B with probability sigmoid((rA − rB)/τ), so a larger τ means noisier labels — with τ near 3 the simulated labeller is close to guessing. Your own clicks are recorded exactly as made, mistakes included. After every comparison the six scores are refit by minimising the same cross-entropy loss the papers use (plus a small L2 penalty so an undefeated swatch cannot run off to infinity). Rank agreement is the share of the 15 swatch pairs that the fit orders the same way as the truth; 50% is what a random ordering scores on average. The left chart shows the fitted scores; the small number under each bar is that swatch's true rank, 1 = darkest. The predicted probability in the readout is the fit's own estimate for the pair currently shown.

Turning choices into scores

The step that makes this work is a piece of statistics. Suppose each output has a hidden quality, a number r, and the probability that a person prefers A to B is the sigmoid of the difference — a curve that gives 50% when the scores are equal and climbs toward certainty as A pulls ahead. This is the Bradley–Terry model, the rule behind chess Elo ratings: a rating gap predicts a win probability. Both papers build their reward model on it.

Fitting it is ordinary supervised learning. A network maps an output to a score. For each comparison, compute the sigmoid of the score difference and nudge the network so the winner's score rises relative to the loser's. Over the whole pile the scores settle into an ordering explaining most choices. Only differences matter, so the scores have no natural zero; Ouyang's team shifted theirs so human demonstrations averaged zero.

Why comparisons? Christiano's group found that in some domains people give consistent comparisons more easily than consistent scores.

Why the policy is kept on a leash

Once the reward model exists, the original model — the policy — is trained by reinforcement learning to score highly under it. But the reward model only knows outputs like those it saw compared. Push the policy far from there and it can find outputs the reward model scores well for no good reason. Christiano's team saw this with a reward model frozen after training: a Pong agent sometimes learned to avoid losing points rather than score them, volleying endlessly. Ouyang's team added a penalty on the KL divergence — how far the tuned model's word choices have drifted from those of the supervised model it started from — at every token, to limit exactly this over-optimisation. The policy may improve, but not wander where the reward model never looked.

In short

People find it easier to compare two outputs than to score one. Bradley–Terry turns a pile of comparisons into a score for every output. A model trained to raise that score, while staying near its starting point, ends up doing what people preferred — without anyone writing down what that was.

Where this comes from

  1. Deep reinforcement learning from human preferences linked only, not reproduced
    Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei · arXiv:1706.03741 · 2017
    arxiv.org/abs/1706.03741
  2. Training language models to follow instructions with human feedback linked only, not reproduced
    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, Ryan Lowe · arXiv:2203.02155 · 2022
    arxiv.org/abs/2203.02155