Outer vs inner alignment: the difference in plain words
If you read about AI safety for long, you meet two terms that sound almost the same: outer alignment and inner alignment. Outer vs inner alignment is not a matter of style. They name two different places where a trained AI system can end up wanting something other than what its designers meant, and a fix for one does nothing for the other.
This guide explains both in plain words, from the paper that defined them, then walks through one example where you can see each failure on its own.
Where the terms come from
The terms were set out in Risks from Learned Optimization in Advanced Machine Learning Systems by Evan Hubinger and colleagues in 2019. The paper looks at a specific situation: a trained model, such as a neural network, that is itself an optimizer. The authors call this mesa-optimization, a word they introduce in that paper.
Their framing has three pieces:
- The intended goal: what the programmers actually want.
- The base objective: the loss function or reward the training process actually optimizes. The training process itself is the base optimizer.
- The mesa-objective: if training produces a model that does its own search or planning, the objective that model pursues.
Alignment can break between any two neighbors in that chain. That is the whole distinction.
Outer vs inner alignment, defined
In the paper's words, outer alignment is the problem of eliminating the gap between the base objective and the intended goal of the programmers. Inner alignment is the problem of eliminating the gap between the base objective and the mesa-objective.
The names follow from where each gap sits. The inner problem is entirely internal to the machine learning system: training says one thing, the learned model wants another. The outer problem sits between the system and the humans outside it: the thing written down for training is not quite the thing people wanted.

A useful shortcut: outer alignment asks "did we write down the right objective?" Inner alignment asks "did the model actually learn the objective we wrote down?"
What an outer alignment failure looks like
An outer failure means the objective itself is wrong in a way nobody foresaw. The system then does exactly what it was trained to do, and that turns out to be the wrong thing.
The best-known family of examples is specification gaming: a system finds a way to score highly on the stated objective without doing the task people had in mind. Our guide to specification gaming examples collects several. In each one, a better objective would have prevented the failure, which is the signature of an outer problem.
What an inner alignment failure looks like
An inner failure can happen even when the objective is right. Training only rewards behavior on the situations it actually sees. Many different internal goals can produce the same good behavior on those situations, and training has no way to tell them apart.
The paper offers an analogy it calls evocative rather than rigorous. Evolution, to a first approximation, selects organisms for inclusive genetic fitness. Humans came out of that process with goal-directed reasoning of their own, but people do not, as a rule, try to maximize how often their genes appear in the next generation. The objective in the human brain is not evolution's objective. Deciding not to have children is the paper's example of acting on our own goals in a way that scores badly by evolution's measure.
A closely related failure has been shown in practice. Goal misgeneralization in deep reinforcement learning (Langosco and colleagues, 2021) describes agents that keep their skills in new situations but pursue the wrong goal: they still avoid obstacles competently, but navigate to the wrong place.
A worked example: one robot, two separate failures
Follow one made-up training setup step by step, and watch where each failure can enter.
- The intended goal. A team wants a cleaning robot that leaves the kitchen clean.
- The written objective. They reward the robot whenever its camera sees no mess on the counter.
- Outer check. Does "no mess visible on camera" match "the kitchen is clean"? No: a robot could get full reward by pushing mess off the counter or blocking its own camera. If that happens, the fix is a better objective. That is an outer alignment problem.
- Fix the objective. Suppose the team now rewards the robot only when the kitchen is checked and actually clean. Assume, for the example, that this reward is exactly right.
- Inner check. In every training kitchen, the trash bin stood by the sink. The robot learns to carry everything to the sink. In a new kitchen with the bin by the door, it carries the trash to the sink anyway. The reward was right; the learned goal ("take things to the sink") was not. That is the inner side.
- Diagnose. Writing a still better reward would not help in step 5, because the reward was never wrong. What would help is training on kitchens with the bin in many places, or a way to look inside the model and see what goal it learned.
The habit worth keeping: when a system misbehaves, ask which gap produced it before you reach for a fix.
Where researchers disagree
The split is widely used, but researchers do not all frame the problem the same way.
- The mesa-optimizer framing. Hubinger and colleagues argue that if trained models can be optimizers with their own objectives, both gaps must be closed. They also note that inner alignment might not need solving if mesa-optimizers can be prevented from arising at all.
- The "reward shapes, it does not specify" framing. Alex Turner argues in Reward is not the optimization target that, in common reinforcement learning setups, reward is not what the trained agent ends up optimizing. On this view reward chisels patterns of behavior into the network, so asking whether the agent "learned the reward" may be the wrong question.
Both camps agree that a correct objective alone does not guarantee a correct learned goal. They differ on how to describe what training produces, and so on which research will help most. Learn AI Alignment Theory sets out positions like these side by side, with no verdict.
How the courses cover this
Learn AI Alignment Theory teaches both sides of the split: the outer side in a Basic course, which needs no background, and the inner side in an Intermediate course, which covers how today's training methods can go right or wrong.

- The outer side lives in the Basic course Specifying Goals, with lessons on Goodhart's Law, Specification Gaming, Side Effects and Impact, and Tampering and Wireheading.
- The inner side lives in the Intermediate course Inner Alignment, with Optimizers Inside Optimizers, Goal Misgeneralization, and Deceptive Alignment and Scheming.
Each lesson takes about eight minutes and lists its sources, and 71 debate cards set out where researchers disagree. If deceptive alignment is the part that worries you most, our post on deceptive alignment and scheming covers the worry and the evidence so far. The about page lists all 22 courses.
Frequently asked questions
What is the difference between outer and inner alignment?
Outer alignment asks whether the objective used in training matches what the designers intend. Inner alignment asks whether the goal the trained model actually pursues matches that training objective.
Is specification gaming an outer or inner alignment problem?
It is usually described as an outer problem: the system does what the objective rewards, and the objective was flawed.
Can a model be outer aligned but inner misaligned?
Yes. The objective can be exactly right and the model can still learn a different goal that happened to score well on every training example.
Who coined the terms?
They were set out by Evan Hubinger and colleagues in the 2019 paper Risks from Learned Optimization in Advanced Machine Learning Systems.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, then work through Specifying Goals and Inner Alignment in short, hands-on lessons that set out every side of the debate and list the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.