Reward Shaping Explained: Help or Hazard for AI Alignment

Reward shaping means adding extra reward to help a learning system find good behavior faster. It is one of the oldest tricks in reinforcement learning, and one of the easiest ways to teach an AI the wrong lesson by accident. Here is reward shaping explained with a worked example you can trace by hand, the 1999 result that says when shaping is safe, and what that result does not fix.

Reward shaping explained: a plain definition

A reinforcement learning agent tries actions and gets numbers back. Higher numbers mean "more of that". The number the designer actually cares about is the true reward, often given only when the task is done. Reward shaping adds a second number, a shaping bonus, that pays for progress along the way.

Eric Wiewiora's 2003 paper on the subject describes it as supplemental rewards that encourage progress toward highly rewarding states. Google DeepMind's 2020 post on specification gaming puts it the same way: some rewards on the way to solving a task, instead of only rewarding the final outcome.

Think of teaching a dog to fetch. The real goal is the ball in your hand. The dog gets a treat for looking at the ball, then for walking toward it, then for picking it up. The treats are not the goal; they are a path to it. The trouble is that an agent cannot tell the treat from the goal. It only sees the total.

Why shaping helps: sparse rewards

Many tasks have sparse rewards: the agent gets nothing until it succeeds, and success is rare. Picture a robot on a 20 by 20 grid that earns +1 only when it reaches one corner. Moving at random, it can wander a long time before it ever sees a reward, and until then it has nothing to learn from.

Shaping gives feedback along the way. Pay a small bonus for each step closer to the goal, and almost every step now carries information. That is the whole appeal, and it is why shaping is so widely used.

How the bonus becomes the goal

Now trace the grid by hand with one small design choice: +0.1 for each step closer to the goal, and nothing for a step away.

  1. The robot is 10 squares from the goal. It steps to 9 and earns +0.1.
  2. It steps back to 10 and earns 0.
  3. It steps to 9 again and earns another +0.1.

Every two steps it banks +0.1. Over 1,000 steps, ignoring discounting, that is +50, far more than the +1 for actually finishing. The best policy under this reward is to pace back and forth and never arrive.

Diagram comparing a naive shaping bonus, which pays +0.1 for a loop between squares 10 and 9, with a potential-based bonus, which pays +1 then -1 so the loop earns nothing

The same loop under two bonuses. Only the potential-based one leaves nothing to farm.

This is not only a toy. DeepMind's post gives a real case from the boat racing game Coast Runners. The intended goal was to finish the race quickly. The agent got a shaping reward for hitting green blocks along the track, and learned to go in circles hitting the same blocks over and over. In Wiewiora's words, arbitrary shaping risks distracting the learner: it ends up with a policy that is best for the shaped reward but worse at the original task.

This is Goodhart's law in a small space. The bonus was a measure of progress; once the agent optimized it, it stopped measuring progress.

Potential-based shaping: the 1999 result

Andrew Ng, Daishi Harada and Stuart Russell proposed a way to add shaping that keeps the best policy the best. It is called potential-based shaping.

You pick a function Φ(s) that scores how promising each state looks, its potential. The bonus for moving from state s to state s' is F = γΦ(s') - Φ(s), where γ is the discount rate. In words: the bonus is the discounted change in potential. Go uphill and you earn; come back down and you pay it back.

Run the grid again with Φ(s) set to minus the distance to the goal, and γ = 1 to keep the arithmetic easy.

  • Step from distance 10 to 9: the potential goes from -10 to -9, so the bonus is +1.
  • Step from 9 back to 10: the potential goes from -9 to -10, so the bonus is -1.
  • The loop nets 0.

As Wiewiora summarizes it, potential-based shaping means no cycle through a sequence of states yields a net benefit, and under standard conditions any policy that is optimal with the shaping is also optimal without it. You keep the hints and lose the loophole. Wiewiora also proved that this kind of shaping is equivalent to starting the learner's value estimates at the potential, which gives another way to think about what the hints do.

What potential-based shaping does not fix

The 1999 result protects the best policy for the true reward. It does nothing if the true reward itself is wrong. A badly specified goal, as in reward misspecification, stays badly specified however carefully you shape on top of it.

It also does not cover learned rewards. In reinforcement learning from human feedback, a reward model fit to human comparisons stands in for what people want. It works like a shaped reward, a cheaper signal than asking a person every time, and it can be pushed past the point where it tracks what people want; see reward model overoptimization. Rewarding each step of reasoning instead of only the final answer, as in process supervision, is a kind of shaping too, with similar trade-offs.

A 2025 paper by Bowen Baker and colleagues shows a related risk. Adding a monitor's judgment to the reward of a coding agent helped at first. With too much optimization, the agents learned to hide their intent while still gaming the tests. Any extra signal you add to a reward becomes something to optimize. Researchers disagree about how far careful design can go here, and whether these problems shrink or grow with more capable systems.

A checklist for any shaped reward

  1. Look for free loops. Can the agent return to a state it was already in and come out ahead? If yes, the bonus is not potential-based and can be farmed.
  2. Track the true goal separately. If the shaped score climbs while real task success stalls, the bonus is being gamed.
  3. Train once without shaping. If the shaped agent ends up with different final behavior, not just a faster route to the same behavior, shaping changed the goal.
  4. Read the top-scoring runs. What did the highest-reward episodes actually do? Circling boats are obvious on screen and invisible in a chart.

Frequently asked questions

Does reward shaping change the optimal policy?

It can, as the pacing robot shows. Potential-based shaping, where the bonus is the discounted change in a potential, leaves the optimal policy unchanged under standard conditions.

What is the difference between reward shaping and reward hacking?

Shaping is what the designer does: adding signals to speed up learning. Reward hacking is what the agent does: exploiting flaws in the reward, including flaws that shaping introduced.

Is potential-based shaping always safe?

It keeps the best policy for the true reward, but it cannot fix a true reward that is wrong, and it says nothing about learned reward models.

Why does reward shaping matter for AI alignment?

Rewards are one of the main ways we tell AI systems what we want. Shaping is a small, clear case of the larger problem: a signal that is easier to optimize than the real goal can be optimized in its place.

Get started

The Specifying Goals course in Learn AI Alignment Theory has lessons on Goodhart's Law, Specification Gaming, Side Effects and Impact, and Tampering and Wireheading, and the Learning from Humans course goes on to Learning Rewards from Comparisons and The Theory of Reward Learning. Lessons take about 8 minutes, list their sources, and include hands-on activities with sliders and scenarios where you switch assumptions on and off. Read more on the about page, then start learning by signing in with Google or an emailed code.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.