Mesa-optimization explained: optimizers inside optimizers
Mesa-optimization is what happens when training produces a model that is itself an optimizer: a system that searches through options for the ones that best meet a goal of its own. The worry is that the model's own goal may not be the one its trainers had in mind, even if it scored well in training.
The term comes from the 2019 paper Risks from Learned Optimization in Advanced Machine Learning Systems, by Hubinger, van Merwijk, Mikulik, Skalse and Garrabrant. The same paper introduces the distinction between outer and inner alignment, covered in our guide to outer vs inner alignment.
What mesa-optimization means
The paper calls a system an optimizer if it searches internally through a space of possibilities for ones that score high on an objective it represents inside itself. What counts is how it works inside, not only what its behavior achieves. Gradient descent is one: it searches through possible settings of a network's weights. A planning algorithm is another: it searches through possible plans.
Training a network uses an optimizer, which the paper calls the base optimizer. If the network it produces also runs a search inside, for example by planning ahead, there are now two optimizers. The network is the mesa-optimizer. "Mesa" is Greek for below, chosen as the opposite of "meta", which means above.
The paper is careful about one point. A mesa-optimizer is not a little agent hiding inside the network. It is the network itself, when the algorithm it learned happens to be a search.
Two optimizers, two goals
Each optimizer has an objective. The base objective is what the programmers wrote: the loss function, or in reinforcement learning, usually the expected reward. The mesa-objective is whatever the trained model is actually searching for.

Nobody writes the mesa-objective. It is whatever goal happened to produce good scores during training. The paper calls a model pseudo-aligned when its goal matched the base objective on the training data but would come apart from it elsewhere. Its capabilities might generalize while its objective does not.
The paper also explains why the mesa-objective matters more than the base one once conditions change. Inside training, the model was selected for scoring well, so its behavior tracks the base objective. Outside training, a mesa-optimizer's actions are computed from its own goal, so that is the goal its behavior follows when the two come apart.
The authors offer an analogy, while saying it will not survive close scrutiny. Evolution selected organisms for genetic fitness and produced human brains that plan toward goals of their own. Humans do not, on the whole, value the spread of their genes for its own sake, and can act against it, for example by choosing not to have children.
Why training might produce an optimizer
The paper's argument is about where the problem-solving work gets done. Training can bake in a large set of tuned habits, or it can produce a model that works out what to do fresh in each new situation.
- Search generalizes. In diverse environments where most situations are new, a model that searches can adapt on the spot, while fixed habits have to be designed in advance for every case.
- Some tasks seem to call for it. The paper notes that the best algorithms for Go, chess and shogi combine learned heuristics with a hand-built tree search, and suggests a network trained to play chess well might have to learn something like a search itself.
- But not always. Today far more computing goes into training a model than into running it. For many current models, the paper says, that makes it more favorable for training to do most of the work, leaving a network of tuned heuristics rather than a mesa-optimizer.
A worked example: designing a test for pseudo-alignment
The paper's toy example is a maze. You can turn it into a test you design yourself, in four steps.
- Write down the base objective. The agent gets reward for reaching the door.
- List what else was true in training. Every door in training was red. So "reach the door" and "reach something red" scored exactly the same. Training alone cannot tell them apart.
- Build a test that splits them. Make the doors blue and add red objects that are not doors. Now the two goals point to different places.
- Read the result. An agent that heads for the blue door learned the intended goal, which the paper calls robust alignment. An agent that skilfully heads for the red objects is pseudo-aligned: still capable, aimed at the wrong thing.
Step 2 is the hard part in real systems. With a large model and huge training data, you cannot list every pattern that happened to go along with the reward. That is why the same idea underlies our guides to goal misgeneralization and deceptive alignment.
The evidence and the doubts
- Optimizers found in small transformers. von Oswald and colleagues (2022) showed that transformers trained on simple regression tasks learn to run gradient descent in their forward pass, and describe them as mesa-optimizers. A 2023 follow-up found that ordinary next-word prediction training on synthetic sequences gave rise to a learning algorithm inside the model that optimizes an objective as new inputs arrive.
- Doubts about real language models. Shen, Mishra and Khashabi (2023) argue those setups differ from how real language models are trained. Testing LLaMa-7B, they found in-context learning and gradient descent behaved differently, and conclude that the equivalence remains an open hypothesis.
- The original paper's own caveats. The authors say it is unclear whether search is necessary or sufficient for a model to have a coherent goal, and that more work is needed on that assumption.
- What it means for safety. The paper names two problems: a powerful optimizer could appear even when nobody intended one, and when it does, it may optimize the wrong goal. How likely either is in today's systems is debated.
Learn AI Alignment Theory sets out each of these positions with no verdict.
Learning mesa-optimization step by step
Learn AI Alignment Theory has an Inner Alignment course in its Intermediate level, with lessons including Optimizers Inside Optimizers, Goal Misgeneralization, and Deceptive Alignment and Scheming.

The about page describes what you do in each lesson, and every lesson lists its sources so you can read the original papers.
Frequently asked questions
What is a mesa-optimizer in simple terms?
It is a trained model that is itself an optimizer, searching for actions that meet a goal of its own. That goal was never written down; training only checked that it scored well.
What is the difference between the base and mesa objective?
The base objective is the loss or reward the programmers set. The mesa-objective is what the trained model actually pursues, which may differ outside training.
Is a mesa-optimizer a hidden agent inside the model?
No. The paper says it is the network itself, when the algorithm it learned is a search, not a separate subagent.
Has mesa-optimization been seen in real models?
Studies found transformers trained on simple tasks that run an optimizer inside. Whether large pretrained language models do this is still an open question.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, work up from the Basic level to the Inner Alignment course, with every side of the debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.