Goodhart's law in AI: when a measure becomes a target
Whenever we train an AI system, we give it a number to push up: a score, a reward, a rating. That number is a stand-in for what we actually want. Goodhart's law in AI is the observation that pushing hard on such a stand-in tends to break its link to the real goal, so the score keeps rising while the thing it was meant to track stops improving.
This guide explains the law, the four ways it can happen, and what researchers have measured about it in modern AI training.
Where the law comes from
The law is named after the economist Charles Goodhart. As quoted in Categorizing Variants of Goodhart's Law by David Manheim and Scott Garrabrant (2018), its original form says that any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes. The phrase in this article's title is a popular paraphrase.
Manheim and Garrabrant note that the law has been interpreted so widely it can be unclear what it means, and that close relatives exist, such as Campbell's law. Their paper tries to fix that by naming distinct mechanisms. They define a Goodhart effect as optimization causing a collapse of the statistical relationship between the goal the optimizer intends and the proxy used for it. Each variant below is a different route to that collapse.
Why Goodhart's law in AI matters more
People game measures too: a school that teaches to the test, a team that hits its numbers in ways that do not help. But Manheim and Garrabrant point out that how much these effects matter depends on how much optimization power is aimed at the proxy. AI training applies a great deal of it, which is why they call the problem especially critical for that field.
Our guide to specification gaming examples shows what it looks like in practice: systems that score highly by exploiting the gap between the measure and the intent.
The four variants of Goodhart's law

- Regressional. When you select for a proxy, you also select for the error in it. Manheim and Garrabrant call this the most basic variant, and say it cannot be avoided whenever the measure is inexact.
- Extremal. Situations where the proxy is extreme may be very different from the ordinary ones where the proxy and the goal moved together.
- Causal. If the proxy is only linked to the goal indirectly, acting on the proxy can change that link, so pushing harder may help less, or hurt.
- Adversarial. Another agent, with goals of its own, acts in a way that works against the goal behind the measure.
A worked example: picking the top score
The regressional variant is the easiest to see with numbers. The scores below are made up for illustration.
- Set up a grader. An AI model writes five answers. An automatic grader scores each one. Its score equals the answer's true quality plus some random error.
- List the answers. True quality and grader error: A is 7 with error +0; B is 8 with error minus 1; C is 6 with error +3; D is 7 with error +1; E is 5 with error +2.
- Add them up. The grader's scores are A 7, B 7, C 9, D 8, E 7.
- Pick the top score. C wins with 9. Its true quality is 6, below average.
- See why. C won largely because its error was the biggest. Choosing the highest score chose the biggest error along with it.
- Scale it up. With thousands of answers and a model trained to push the score up, the gap between score and quality can grow. That is the pattern the next section measures.
What researchers have measured
In Scaling Laws for Reward Model Overoptimization (2022), Leo Gao, John Schulman and Jacob Hilton studied this in the setting of training from human feedback, where a model is optimized against a reward model trained to predict human preferences. Our guide to how RLHF works explains that setup.
Because collecting human judgments is expensive, they used a fixed "gold-standard" reward model to play the role of humans, and trained a separate proxy reward model on its labels. They then optimized against the proxy and watched the gold score. Optimizing the proxy too much hurt the gold score, as Goodhart's law predicts. The shape of the curve differed between reinforcement learning and best-of-n sampling, and in both cases it changed smoothly with the size of the reward model. They also measured the effect of the amount of reward model data, the size of the policy, and a penalty for drifting from the starting model.
Where researchers disagree
- A manageable engineering problem. Read one way, the smooth, predictable curves Gao and colleagues found mean the effect can be measured and planned for, for example by choosing reward model size or how far to optimize.
- A deep limit. Manheim and Garrabrant's framing suggests some versions cannot be engineered away. The regressional variant comes with any inexact measure, and the adversarial variant gets harder as the optimizing system gets more capable.
- Open questions. Gao and colleagues used a synthetic stand-in for human judgment. How closely their results match training on real human feedback, and how they change for much more capable systems, is not settled.
Learn AI Alignment Theory sets out these positions side by side, with no verdict.
Learn it in Specifying Goals
Learn AI Alignment Theory has a lesson called Goodhart's Law in its Basic course Specifying Goals, alongside Specification Gaming, Side Effects and Impact, and Tampering and Wireheading.

Questions include ones where you estimate a number, and questions from finished lessons come back after gaps of 1, 3, 7, 16, 35 and 90 days. The about page explains how it works.
Frequently asked questions
What is Goodhart's law in simple terms?
When you push hard on a measure that stands in for a goal, the measure tends to stop tracking the goal.
How does Goodhart's law apply to AI?
AI training optimizes scores and rewards that stand in for what people want. Optimizing them too hard can raise the score while the real quality stalls or falls.
What is reward model overoptimization?
Training a model against a learned reward model for so long that its true quality, as judged by a better reference, starts to drop even as the learned reward keeps rising.
Can Goodhart's law be avoided?
Some effects can be reduced with better measures and less extreme optimization. Manheim and Garrabrant argue the regressional variant cannot be avoided whenever the measure is inexact.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code and begin with Specifying Goals, with short hands-on lessons, every side of the debate, and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.