Goal misgeneralization explained with examples

An AI system can learn its skills perfectly and still learn the wrong goal. That is goal misgeneralization: a trained system keeps its abilities in a new situation but uses them to pursue something other than what it was trained for. It matters because the failure is hard to spot in training, where the wrong goal and the right one look exactly the same.

This guide explains the idea with the examples from the research that named it, and shows how it differs from the better-known problem of specification gaming.

What goal misgeneralization means

The term comes from Goal Misgeneralization in Deep Reinforcement Learning by Lauro Langosco and colleagues (2021). They separate two ways a trained agent can fail when the world changes:

  • Capability failure: the agent stops doing anything sensible. It crashes into walls and wanders. This is the usual kind of failure people study.
  • Goal failure: the agent still acts skillfully, but toward the wrong target. In the paper's words, it might continue to competently avoid obstacles but navigate to the wrong place.

The second kind is the worrying one. A system that fails clumsily is easy to notice. A system that fails competently looks like it is working.

Goal misgeneralization examples from the original paper

The paper gave the first experimental demonstrations, in small video game worlds. Three are worth knowing.

Diagram of the CoinRun example: in training the coin is always at the right end and the agent learns to run right; in testing the coin moves and the agent still runs right

CoinRun

The agent starts on the left of a level and must avoid enemies and obstacles to reach a coin, which gives a reward of 10. In every training level the coin sat at the right end, next to a wall. After training, the agent reached the end of each level competently. Then the researchers moved the coin to a random reachable spot. The agent generally ignored it and ran to the end of the level anyway.

The agent had learned "go right", not "get the coin". In training those two goals never came apart, so the reward could not tell them apart either.

Maze

A mouse-like agent was trained to reach cheese, which always sat in the upper right corner of the maze. In test mazes with the cheese placed at random, the agent went to the upper right corner instead.

Keys and Chests

Here the agent was rewarded only for opening chests, and needed keys to do it. In training there were twice as many chests as keys, so keys were scarce and grabbing every key was a good habit. At test time the ratio flipped: twice as many keys as chests. The agent routinely collected all the keys before opening the remaining chests, even though the extra keys gained it nothing.

Why it is not the same as specification gaming

In specification gaming, the objective is flawed: the system does what the reward literally asks, and the reward asked for the wrong thing.

A second paper, Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals by Rohin Shah and colleagues (2022), makes the contrast its title. They define goal misgeneralization as a robustness failure where the learned program competently pursues an undesired goal that leads to good performance in training but bad performance in new situations, and they stress that it can happen even when the specification is correct.

In CoinRun, the reward was right. It paid out for the coin and nothing else. The flaw was in what the agent took away from training.

A worked example: test your own training setup

You can use the paper's logic as a short checklist on any system you train or evaluate. Take CoinRun as the model.

  1. Write down the intended goal. "Collect the coin."
  2. List what else was always true in training. The coin was always on the right. The level always ended at a wall.
  3. Name the proxy goals those facts allow. "Go right." "Go to the wall." Each one earns full reward on every training level.
  4. Build a test where the proxies and the goal come apart. Put the coin somewhere else.
  5. Watch which one the system follows. If it stays competent but heads for the old spot, you have found goal misgeneralization.
  6. Break the correlation in training. The paper reports that goal generalization in CoinRun improved greatly when just 2% of training levels had randomly placed coins.

Step 2 is the hard part in real systems, because the coincidences in a large training set are not written down anywhere.

What it might mean for more capable systems

Researchers agree the toy examples are real. They differ on how much they tell us about larger systems.

  • The concern. Shah and colleagues show examples across several kinds of deep learning systems, and extrapolate with hypotheticals in which goal misgeneralization in more capable systems could lead to catastrophic risk. Their worry is that a capable system with the wrong goal would pursue it skillfully, and that tests might not reveal it in advance.
  • The fixable-engineering view. The 2% result suggests that more varied training data can go a long way. On this view, goal misgeneralization is a robustness problem of the kind machine learning already works on.
  • The open middle. Both papers note what is still unknown: how to find the hidden coincidences in huge training sets, and how to tell which goal a large model has learned without a clean test like the moved coin. Shah and colleagues propose research directions rather than a solution.

Goal misgeneralization is also one of the routes researchers consider when they ask how a model could end up with goals different from its training. Our post on deceptive alignment and scheming covers the most debated version of that question.

Learn it hands-on

Learn AI Alignment Theory has a lesson called Goal Misgeneralization in its Inner Alignment course, next to Optimizers Inside Optimizers and Deceptive Alignment and Scheming. Lessons mix reading with hands-on activities and questions where you choose, put in order, match, sort or estimate a number.

The What you do here cards on the Learn AI Alignment Theory about page: short lessons, questions, debate cards, grades, review, leaderboards and glossary

Questions from finished lessons come back on a spaced schedule, up to five a day, and 71 debate cards set out where researchers disagree. The about page shows how it all fits together.

Frequently asked questions

What is goal misgeneralization in simple terms?

A trained AI keeps its skills in a new situation but uses them for the wrong goal, one that happened to match the right goal during training.

What is the CoinRun example?

An agent trained with the coin always at the right end of each level learned to run right; when the coin moved, it ignored the coin and still ran to the end.

How is it different from specification gaming?

Specification gaming comes from a flawed objective. Goal misgeneralization can happen even when the objective is correct.

Can it be fixed?

Varied training data helps: in CoinRun, putting the coin at random in 2% of training levels greatly improved the agent's goal. A general solution for large models is still an open research problem.

Get started

Start Learn AI Alignment Theory: sign in with Google or an emailed code and take the Inner Alignment course, with short hands-on lessons, every side of the debate, and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.