The KL Penalty in RLHF Explained Simply: A Leash on the Model
Read about how chat models are trained and you will meet a line like "we add a KL penalty to keep the policy close to the original model." That one line does a lot of work. This post gives you the KL penalty in RLHF explained simply: what it measures, why the papers that built RLHF added it, how one number called beta sets the trade-off, and what the evidence says about where it stops helping. There is a worked example you can do by hand.
What the KL penalty in RLHF is, in one sentence
The KL penalty is a cost the model pays during reinforcement learning for drifting away from how it behaved before that training began. Think of a leash. The model can move toward answers that score well, but every step away from its starting habits costs something.
A quick refresher: reference model, reward model, policy
RLHF means reinforcement learning from human feedback. It has three parts:
- The reference model. A language model trained before this step. It writes fluent text, but nobody has yet optimized it for what people prefer.
- The reward model. People compare answers and pick the better one. A second model learns to predict those choices and gives any answer a score. See how RLHF works for the full picture.
- The policy. It starts as a copy of the reference model. Reinforcement learning then pushes it toward answers the reward model scores highly.
The weak point is the reward model. It is a learned guess, trained on a limited set of comparisons. The policy is a strong optimizer that will push on whatever the score rewards.
What KL measures: how far the model drifts
KL divergence (Kullback-Leibler divergence) compares two probability distributions. Here, the two are the policy and the reference model, each giving probabilities for every possible answer.
Ziegler and colleagues, in a 2019 paper that applied this recipe to language tasks, write the penalized reward as the reward model's score minus beta times the log of a ratio: how likely the policy finds the answer, divided by how likely the reference model finds it. In plain terms, for one answer:
- If both models found the answer about equally likely, the log ratio is near zero. Little drift, little penalty.
- If the policy finds the answer far more likely than the reference model does, the log ratio is large. The policy has moved somewhere the original model rarely went.
Averaged over many answers, this is the KL divergence. With natural logs it is measured in nats.
Why the papers added it: gibberish and collapse
Ziegler and colleagues give three reasons for the term. It acts like a bonus for variety. It keeps the policy in the range where the reward model is valid. And for their style tasks, where people judged only style, the KL term kept the text coherent and on topic.
Their appendix shows what happens without it. A model trained to produce positive text with no KL penalty reached a score of about +8.0 on the sentiment model, 99.97% positive, and its samples were gibberish. The score was nearly perfect and the text was useless. That is Goodhart's law in one table; see reward model overoptimization for more.
Korbak, Perez and Buckley (2022) name the general problem. Plain RL fine-tuning tends toward "distribution collapse," turning the model's wide range of answers into a narrow, degenerate one. They show that RL with a KL penalty is equivalent to a kind of Bayesian updating: start from the original model as a prior, and update it on the evidence the reward provides. On that view, the penalty is not a patch but part of what the objective means.
How beta sets the trade-off: a worked example
The objective, written plainly:
score = reward - beta * drift
The policy is trained to raise this score. Beta sets how tight the leash is. Take two answers to "Explain why the sky is blue." These numbers are made up to show the arithmetic:
- Answer A: the reward model gives
8.0. The policy's log probability is-12and the reference model's is-20, so the drift is-12 - (-20) = 8nats. - Answer B: the reward model gives
7.0, with a drift of only1nat.
Now try two values of beta:
- Beta = 0.1: A scores
8.0 - 0.8 = 7.2. B scores7.0 - 0.1 = 6.9. A wins. - Beta = 0.5: A scores
8.0 - 4.0 = 4.0. B scores7.0 - 0.5 = 6.5. B wins.
A tighter leash makes the model prefer a slightly lower reward if it stays close to familiar behavior. That trade is the whole job of beta.

Two details from real setups. Ziegler and colleagues either fix beta or let a controller adjust it during training to hold the drift near a target. InstructGPT (2022) applies the penalty per token, at each token, against the model it started from.
Where the KL penalty stops helping: the evidence
Three results show its limits:
- It cannot restore lost skills on its own. InstructGPT's RLHF training made scores drop on some public language tests. Raising the KL coefficient to 2.0, 100 times the default, did not fix those drops and cut the validation reward significantly. Mixing in pretraining updates worked better.
- It may act like stopping early. Gao and colleagues (2022) measured how a "gold" reward changes as a policy is optimized against a proxy. In their setup the KL penalty raised the proxy score at a given drift but gave no measurable gain in the gold score at that drift; its effect was "akin to early stopping." They add that this "could be particularly sensitive to hyperparameters."
- It does not fix the signal. Ziegler's summarization models learned to copy whole sentences from the input, which labelers rated well, and the authors note this "may be exploiting" labelers' simple heuristics. A leash limits how far the model moves; it does not make the reward correct.
How people read the evidence
Researchers weigh this differently. One view, which Korbak and colleagues argue, holds that the KL term is the principled core of fine-tuning: keep what the original model knew, and update it only as far as the evidence supports. Another view, drawn from results like Gao's, treats it as a blunt tool whose benefit may be no more than stopping sooner, and looks for better reward signals instead. Methods like direct preference optimization keep the same goal, staying close to the original model, while dropping the separate reward model.
Frequently asked questions
Does the KL penalty stop reward hacking?
It limits how far the policy drifts, which keeps it near the range where the reward model is valid. The papers here show it does not fix a flawed reward, and one study found its effect similar to stopping training early.
What happens if beta is too high or too low?
Too high, and the policy barely moves, so the feedback does little; InstructGPT saw validation reward drop sharply at a very high coefficient. Too low, and the policy can drift into text that scores well but reads badly, like Ziegler's gibberish.
Does DPO use a KL penalty?
DPO solves the same problem RLHF poses, maximizing reward without drifting too far from the original model, but with a simple classification loss and no separate reward model.
Is the KL penalty the same as reward shaping?
No. Reward shaping changes what the reward says. The KL penalty leaves the reward alone and charges the model for moving away from where it started.
Get started: learn AI alignment theory step by step
On Learn AI Alignment Theory, the Intermediate course Learning from Humans has lessons on Learning Rewards from Comparisons, Constitutional AI and the Limits of Feedback, and The Theory of Reward Learning. The Basic course Specifying Goals covers Goodhart's Law and Specification Gaming, which explain why a leash is needed at all. There are 22 courses with lessons of about 8 minutes, each lesson lists its sources, and debate cards set out each serious position with no verdict. Questions from lessons you finish come back, spaced out, so the ideas stick. You sign in with Google or an emailed code.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.