Direct preference optimization explained simply

Direct preference optimization, or DPO, is a way to teach a language model which answers people prefer without building a separate reward model and without reinforcement learning. You give it pairs of answers to the same prompt, labeled with the one people liked better, and it trains the model with one simple loss. It was introduced in 2023 and is now widely used.

This guide explains what DPO removes from the usual pipeline, walks one preference pair through it, and sets out where researchers disagree about how good it is.

The problem DPO was built to solve

The usual way to tune a chat model on human preferences is reinforcement learning from human feedback, or RLHF. Our guide to how RLHF works covers it in full. In short, it has three stages:

  1. Fine-tune the base model on good example answers.
  2. Train a reward model, a second network that learns to score answers the way the preference labels do.
  3. Run reinforcement learning (usually an algorithm called PPO). The model writes new answers during training, the reward model scores them, and a penalty keeps the model from drifting too far from where it started.

In Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov, Archit Sharma, Eric Mitchell and colleagues call this process "complex and often unstable". Their full text points out that it means training several models and sampling from the model in the loop of training, which costs a lot of compute.

How direct preference optimization works

The paper's key step is a piece of math. Under a standard model of how people choose between two answers (the Bradley-Terry model), the best policy for the RLHF objective can be written down directly in terms of the reward. Turn that around, and the reward can be written in terms of the policy itself. So the model you are training already defines a reward, and you never need to train one separately. That is the "secretly a reward model" in the title.

The result is one loss you can minimize on a fixed dataset of preference pairs. The authors describe DPO as implicitly optimizing the same objective as existing RLHF methods: get high reward while staying close to the starting model. Two things make it work:

  • A frozen reference model. A copy of the model after stage 1 is kept fixed. Every probability is measured against it, so the model is rewarded for moving toward preferred answers relative to where it began, not for raw probability.
  • A setting called beta. It controls how far the model may move from the reference. A small beta allows bigger moves; a large beta keeps the model close.
Diagram comparing RLHF's three stages (fine-tuning, reward model, reinforcement learning) with DPO's two stages (fine-tuning, then one classification loss on preference pairs measured against a frozen reference model)

A worked example: one preference pair through DPO

Follow a single training example step by step. The numbers below are made up to show the logic; they do not come from any paper.

  1. The data. The prompt is "Explain photosynthesis to a ten-year-old." Answer A uses a simple kitchen comparison. Answer B is accurate but full of jargon. A person preferred A.
  2. Ask both models. Measure how likely each answer is under the model being trained and under the frozen reference. Say that, so far, training has made A twice as likely as the reference finds it, and B exactly as likely.
  3. Compute the margin. DPO looks at the gap: how much more the model has raised A than B, relative to the reference, scaled by beta. Here the gap is positive, so the model is already leaning the right way.
  4. Turn the margin into a loss. A sigmoid squashes the margin into a probability that A is preferred, and the loss is low when that probability is high. This is ordinary binary classification, the kind used to tell spam from not spam.
  5. Update with a weight. The gradient pushes A's probability up and B's down. The paper's important detail is that each example is weighted by how wrongly the model's implicit reward currently orders the pair. Pairs the model already gets right get small updates; pairs it gets backwards get large ones. The authors report that without this weight, the model can degenerate.

Repeat that over thousands of pairs and you have DPO. Notice what never happened: no reward model was trained, and the model never wrote a new answer during fine-tuning.

What the original paper found

The authors tested DPO on three tasks: steering movie reviews toward positive sentiment, summarizing Reddit posts, and single-turn dialogue on Anthropic's helpful and harmless dataset. They report that DPO beat PPO-based RLHF at controlling sentiment, and matched or improved response quality in summarization and dialogue, while being "substantially simpler to implement and train".

For summaries and dialogue, quality was judged by GPT-4 comparing answers. The authors ran a human study to check this and found humans agree with GPT-4 about as often as humans agree with each other.

The method spread quickly. Zephyr (Tunstall and colleagues, October 2023) used a distilled form of DPO on answers ranked by a larger teacher model instead of people, needing only a few hours of training and no human labels. A 2024 paper on SimPO calls DPO "a widely used offline preference optimization algorithm" and proposes a variant that drops the reference model.

Where researchers disagree about DPO

Whether DPO is as good as RLHF, or just easier, is an open question. Here are the serious positions.

  • Simpler and as good. The original paper's results support this view: the same objective, fewer moving parts, stable training, and matching or better quality on its tasks.
  • PPO still wins when tuned well. In Is DPO Superior to PPO for LLM Alignment? (April 2024), Shusheng Xu and colleagues argue that DPO "might find biased solutions that exploit out-of-distribution responses", and show its results depend heavily on how far the model's own outputs are from the preference data. In their benchmarks, from dialogue to code competitions, a carefully tuned PPO beat DPO in every case.
  • It removes one assumption but keeps another. Mohammad Gheshlaghi Azar and colleagues (October 2023) note that RLHF assumes both that pairwise preferences can be turned into single reward scores and that a reward model generalizes to new outputs. DPO avoids the second but "still heavily relies on the first". They propose an alternative, IPO, and show it beats DPO on some illustrative examples.
  • No reward model does not mean no overoptimization. You might expect that dropping the reward model ends reward model overoptimization. A 2024 study by Rafailov and colleagues found that DPO-style methods "still commonly deteriorate from over-optimization", with patterns similar to classic RLHF, often before a single pass through the dataset is complete.

Each side has evidence, and the results depend on the task, the data and how much tuning each method gets.

What DPO does not change for alignment

DPO changes how preferences are turned into training. It does not change where the preferences come from. If the people labeling pairs reward answers that sound confident or agree with them, DPO will learn that too, just as RLHF would; our post on AI sycophancy covers what that looks like. Methods like Constitutional AI change the labels instead, using written principles and AI feedback, and can be combined with either training method.

Frequently asked questions

Is DPO a type of reinforcement learning?

No. DPO trains with a classification loss on a fixed dataset of preference pairs, and its authors describe it as RL-free. It aims at the same objective as RLHF but reaches it without a reinforcement learning loop.

Does DPO need a reward model?

Not a separate one. The paper shows the model being trained defines a reward implicitly, measured against a frozen reference model, so no second network is trained to score answers.

Is DPO better than PPO?

It depends on who you ask and on the task. The original paper found DPO matched or beat PPO on its tasks, while a 2024 study found a carefully tuned PPO beat DPO in all of its benchmarks.

What does beta do in DPO?

Beta controls how far the trained model may move from the reference model. A larger beta keeps the model closer to where it started.

Get started

DPO is one answer to a bigger question: how do you turn what people prefer into what a model does? In Learn AI Alignment Theory, the Intermediate course Learning from Humans works through it in lessons such as Learning Rewards from Comparisons and Constitutional AI and the Limits of Feedback. Every lesson lists its sources, and the debate cards set out each side with no verdict. You sign in with Google or an emailed code.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.