Reward model overoptimization explained: Goodhart's law in RLHF
Reward model overoptimization is what happens when you train an AI model too hard against a learned reward model: the reward model's score keeps rising, but the quality you actually wanted levels off and then falls. It is Goodhart's law showing up inside the standard way chat models are trained. It is also a problem researchers have measured closely enough to fit curves to.
This guide covers where the problem comes from, what the measurements show, an exercise that makes the shape easy to see, the fixes researchers have tested, and what is still open.
What a reward model is, and why it can be overoptimized
In reinforcement learning from human feedback, people compare pairs of answers and pick the better one. A separate model, the reward model, learns to predict those picks. The main model is then trained to produce answers the reward model scores highly. Our guide to how RLHF works walks through the loop.
The catch is that a reward model is a proxy. It learned from a limited set of comparisons, so it has blind spots. Light training against it moves the main model toward answers people really do prefer. Heavy training starts to find the places where the reward model is wrong, and exploits them. Researchers call that overoptimization.
An early sighting in summarization
In Learning to summarize from human feedback (2020), Nisan Stiennon and colleagues trained a reward model on human comparisons of summaries of Reddit posts, then trained a summarizer against it. Their results showed the pattern directly. Optimizing against the reward model first improved the summaries, but then it overfit. As they pushed further, real human preferences fell away from what the reward model predicted, and eventually the reward model became anti-correlated with what people preferred.
The overoptimized summaries, they wrote, were long, low quality and full of odd habits, though they still caught the rough gist of the post. They also noted that the same thing happens when a summarizer is trained against ROUGE, an older automatic score. Their advice was to pay closer attention to how the training signal affects the behaviour you actually want.
Measuring reward model overoptimization
Asking people to label large numbers of answers at every step is expensive, so Leo Gao, John Schulman and Jacob Hilton used a stand-in. In Scaling Laws for Reward Model Overoptimization (2022), a large 6 billion parameter "gold" reward model played the role of the humans. It labelled 100,000 comparisons, and smaller proxy reward models, from 3 million to 3 billion parameters, learned from those labels. The researchers then trained against each proxy and checked the gold score along the way.
To measure how hard they were optimizing, they used how far the trained model had drifted from where it started, measured with a statistic called KL divergence. Their main findings:
- Rise, then fall. The gold score first went up and later came down, while the proxy score kept climbing.
- Predictable curves. The shape followed a simple formula, with one form for best-of-n sampling (pick the best of n answers) and another for reinforcement learning.
- Bigger and better-fed reward models help. Larger proxy reward models reached better gold scores, and more training data gave better gold scores and less overoptimization. Below about 2,000 comparisons, the reward models were barely better than chance.
- A bigger main model does not fix it. Larger policies started better but showed very similar amounts of overoptimization.
- A KL penalty was not a cure. Penalizing drift raised the proxy score for a given distance but did not measurably improve the gold score at that distance. The authors note this result could be sensitive to settings.

The authors are clear about limits. A gold reward model is not a person, and their synthetic setting might not transfer to the real world. Their setup also leaves out a second gap: the one between the labels and what people actually intend.
A worked example: watch a proxy get overoptimized
You can see the shape with a pen and ten short summaries of one news story. It takes about twenty minutes, and you play both the reward model and the human.
- Write the proxy. Make a scoring rule: one point per word from a list of five key terms in the story, plus one point per sentence. That is your reward model. It roughly tracks good summaries, because good ones mention the key terms and cover more ground.
- Write the candidates. Write ten summaries of different quality. Make a few honest and tight, and a few long and stuffed with the key terms.
- Do best-of-n. Shuffle them. Take the first one (best of 1). Then score the first 3 and keep the top one (best of 3). Then score all 10 and keep the top one (best of 10).
- Judge as a human. Now read the three winners and rate each out of 10 for how well it actually summarizes the story.
- Compare. Your proxy score rises from best of 1 to best of 10, by design. Your own rating often rises and then drops, because the strongest search found the padded, keyword-stuffed summary your rule loves.
That is overoptimization in miniature. Real systems search far bigger spaces, so they find much stranger gaps in the proxy than you can write by hand.
Fixes researchers have tested
- Reward model ensembles. Train several reward models and be cautious where they disagree. Coste and colleagues (2023), in a setup like Gao's with 25% label noise added, found this practically eliminated overoptimization for best-of-n and improved performance by up to 70%. For reinforcement learning it always reduced overoptimization, and with a small KL penalty it prevented it at no performance cost in their tests.
- Ensembles with different starting points. Eisenstein and colleagues (2023) found ensembles help, and help more when the reward models differ in pretraining, not only in fine-tuning. But they did not eliminate reward hacking, because all the models shared similar error patterns.
- Constraints per reward. It is increasingly common to combine several reward models. Moskovitz and colleagues (2023) proposed keeping each one within the range where it is still a useful proxy, using constrained reinforcement learning.
Why it matters for alignment, and the debate
Overoptimization can show up as specific bad habits. One study of preference models found that optimizing against them sometimes traded truthfulness for sycophancy, telling users what they wanted to hear. It is also close kin to specification gaming.
Researchers read the evidence in different ways:
- A measurable engineering problem. The curves are predictable, and ensembles plus a small KL penalty prevented overoptimization in one study. On this view, good practice can keep it under control.
- Fixes reduce it but leave a residue. When every reward model makes the same kind of mistake, averaging them does not help, as Eisenstein and colleagues found.
- The measured gap is the smaller one. Gao and colleagues stress that their setup leaves out the gap between labels and real human intent. They write that studying this effect could help build theory that may be critical for avoiding dangerous misalignment of future AI systems.
Learning reward models in depth
Learn AI Alignment Theory covers this ground in its Intermediate course Learning from Humans, with lessons such as Learning Rewards from Comparisons and The Theory of Reward Learning. The Basic course Specifying Goals has a lesson on Goodhart's Law. As the about page says, every lesson lists its sources, and debate cards give each serious position with no verdict.
Frequently asked questions
What is reward model overoptimization?
It is when training against a learned reward model pushes its score up while real quality levels off and then falls. The model has found the reward model's mistakes instead of getting better.
Is reward model overoptimization the same as reward hacking?
They overlap. Reward hacking is the general idea of exploiting errors in a reward model to get a high score; overoptimization names the measured pattern where more optimization against a learned reward model makes true performance worse.
Does a KL penalty prevent overoptimization?
Not on its own in Gao and colleagues' tests, though they say the result may depend on settings. Coste and colleagues found a small KL penalty worked well when combined with reward model ensembles.
Do bigger reward models fix the problem?
They help. Larger reward models and more data gave better true scores, but the rise and fall pattern remained.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, take the Learning from Humans course to see how reward models are built and where they break, with every side of the debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.