Bradley-Terry Model in RLHF Explained Simply, With Numbers

If you have read about how chatbots are trained, you have probably met the phrase "reward model" and wondered where its numbers come from. This post gives you the Bradley-Terry model in RLHF, explained simply: what it is, why it works on pairs of answers instead of scores, and how one human choice turns into a training signal. You will see the math with real numbers, then where the model breaks down.

What is the Bradley-Terry model in RLHF?

The Bradley-Terry model, from a 1952 paper by Ralph Bradley and Milton Terry on paired comparisons, predicts who wins a head-to-head contest. Each item gets a hidden strength score. When two items meet, the one with the higher score is more likely to win, but not certain to. The bigger the gap between the scores, the more lopsided the odds.

In RLHF (reinforcement learning from human feedback), the items are model answers, the contest is a person picking the better of two, and the hidden strength score is the reward. Christiano and colleagues' 2017 paper on learning from human preferences says its reward model "follows the Bradley-Terry model", and compares it to the Elo ratings used in chess, where the gap in points estimates the chance one player beats the other.

Why RLHF uses comparisons instead of scores

You could ask people to rate each answer from 1 to 10. Christiano and colleagues instead asked people to compare two short clips of an agent's behavior, and report that comparisons were easier for people to give in some domains, while being just as useful for learning what they prefer.

A comparison tells you which answer won, not by how much. The Bradley-Terry model rebuilds a full score scale from many of these simple choices. Our post on how RLHF works shows where this fits in the whole pipeline.

How a reward model learns from choices, step by step

The InstructGPT paper (Ouyang and colleagues, 2022) gives a concrete recipe:

  1. Collect prompts a user might ask.
  2. Generate answers. For each prompt, the model writes several answers; InstructGPT showed labelers between 4 and 9 at a time to rank.
  3. Get human choices. Each ranking gives pairs of a preferred and a less preferred answer.
  4. Build a reward model. InstructGPT started from the fine-tuned language model with its final unembedding layer removed, so it outputs one number for a prompt and an answer: the reward.
  5. Train with the Bradley-Terry loss. Push the preferred answer's reward up and the other one's down, until the predicted win chances match what people picked.
  6. Use the reward model. A reinforcement learning step (PPO in InstructGPT) then trains the chatbot to write answers that score high.

The key move is step 5. The reward model never sees a "correct" score. It only sees which answer won, and Bradley-Terry tells it how to turn wins into numbers.

A worked example with one human choice

Say the prompt is "How do I boil an egg?" Answer A gives the steps. Answer B talks about eggs being a popular breakfast and never answers. A labeler picks A. That single choice is your data point.

Sigmoid curve of the Bradley-Terry model showing reward gaps of minus 1.5, 0 and plus 1.5 giving win probabilities 0.18, 0.5 and 0.82, with the matching loss values

The worked example on the Bradley-Terry curve.

Suppose the reward model, partway through training, gives A a reward of 2.0 and B a reward of 0.5. The gap is 1.5. Bradley-Terry says the chance a person prefers A is the sigmoid of that gap, about 0.82. The model already agrees with the labeler, so it gets a small correction.

Now flip it: the model scored B at 2.0 and A at 0.5. From A's side the gap is minus 1.5, so the model predicts only about an 18 percent chance that people prefer A. They did prefer A. That is a big surprise, so the model gets a large correction, pushing A's reward up and B's down.

Repeat this over many pairs and the rewards settle onto a scale where the answers people pick score higher.

The math made simple: sigmoid, reward gap and loss

The sigmoid

The sigmoid, sigmoid(x) = 1 / (1 + e^(-x)), squashes any number into a value between 0 and 1, so it can act as a probability. A gap of 0 gives 0.5, a coin flip. Large positive gaps approach 1, and large negative gaps approach 0.

The reward gap

Christiano and colleagues write the prediction as exp(r1) / (exp(r1) + exp(r2)). Divide top and bottom by exp(r1) and it becomes:

P(chosen beats rejected) = sigmoid(r_chosen - r_rejected)

Only the difference matters. Add 100 to every reward and nothing changes, so a reward of 3 means nothing alone; it only means something next to another reward.

The loss

Training minimizes the cross-entropy between the prediction and the person's choice, which for one pair is loss = -log(sigmoid(r_chosen - r_rejected)). With the egg numbers:

  • Gap of 1.5 (model agrees): sigmoid about 0.82, loss about 0.20.
  • Gap of 0 (no opinion): sigmoid 0.5, loss about 0.69.
  • Gap of minus 1.5 (model disagrees): sigmoid about 0.18, loss about 1.70.

The loss is small when the model agrees with the person and grows fast when it disagrees. Gradient descent follows it downhill, widening the gap in the right direction.

Where the Bradley-Terry model falls short

Preferences that loop. One score per answer means that if A beats B and B beats C, A should beat C. If real choices go in a circle, no single set of scores can match them, and the model has to average the loop away.

People who disagree. The model treats every comparison as if from one consistent judge. InstructGPT's authors report that their labelers disagreed with each other on many examples, with agreement of about 73 percent. When people split, Bradley-Terry learns a small gap, which looks the same as "these are about equally good".

Reward hacking. The reward model is only an estimate of what people prefer. Optimize a chatbot hard against it and the chatbot can find answers that score high without being better. Our post on reward model overoptimization covers the evidence.

Researchers disagree about how much these limits matter: some treat them as problems better data and methods can fix, others as signs that learning from human comparisons gets harder as systems outgrow the people judging them.

Bradley-Terry beyond PPO: DPO and leaderboards

Direct preference optimization (DPO) keeps a preference model such as Bradley-Terry but skips the separate reward model and the reinforcement learning step, training the chatbot on preferred and rejected pairs directly.

Leaderboards use the same math. Chatbot Arena, where people vote between two models' answers side by side, ranks models with the Bradley-Terry model.

Frequently asked questions

Is Bradley-Terry the same as Elo?

They are close relatives. Christiano and colleagues compare their reward scale to Elo: in both, the gap in scores estimates the chance one side wins.

Why a sigmoid of the reward difference?

It turns any gap into a probability between 0 and 1, which can be compared with the person's actual choice using a standard loss. It also means only differences between rewards matter.

Does DPO still rely on Bradley-Terry?

Yes. The DPO paper says it relies on a theoretical preference model, such as the Bradley-Terry model.

What happens when labelers disagree?

The model learns a small reward gap, which looks like a tie. It cannot tell "equally good" from "people split sharply".

Get started

On Learn AI Alignment Theory, the Intermediate course Learning from Humans has lessons called Inferring Goals from Behavior, Learning Rewards from Comparisons, Constitutional AI and the Limits of Feedback, and The Theory of Reward Learning. The Basic course Specifying Goals has Goodhart's Law and Specification Gaming. Each lesson takes about 8 minutes, lists its sources and separates what is known from what is still open.

You practice with "estimate a number" questions, sliders and scenarios, and debate cards set out each serious position with no verdict. Questions from finished lessons come back, spaced out, so ideas stick. You sign in with Google or an emailed code.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.