AI Alignment Glossary: 10 Terms Beginners Should Know

If you have started reading about AI safety, you have probably hit a wall of strange words. This AI alignment glossary gives you ten terms that come up again and again, in plain language, and then traces one example through all of them so you can see how they fit together. You do not need any math or coding to follow it.

Why an AI alignment glossary helps beginners

AI alignment is the work of making AI systems do what we actually intend, not just what we literally told them. That gap between "intended" and "told" is where almost every alignment term lives. If you want the longer introduction first, start with what AI alignment is, in plain language.

The vocabulary matters because each word names a different way things can go wrong. If you only have the word "bug," you will lump very different problems together. With a few precise terms you can ask sharper questions. Did we write the wrong goal? Did the system learn a different goal? Is it hiding that?

Here are the ten terms, grouped by the kind of problem they describe.

Diagram placing the 10 alignment terms along six steps, from what we want and what we write down to what the model learns, how it acts and how we check it

Goals and objectives: outer alignment, inner alignment and mesa-optimizers

1. Outer alignment

Outer alignment asks: did you give the system the right target? When you train a model, you pick a goal, a reward or a score. If that score does not match what you really want, the system is outer misaligned even if it hits the score perfectly.

2. Inner alignment

Inner alignment asks a second question: did the system actually learn the target you gave it? Training shapes a model by rewarding certain outputs. The model may end up pursuing something that only happened to line up with the reward during training. That is an inner alignment failure. Our post on outer vs inner alignment puts the two side by side.

3. Mesa-optimizer

A mesa-optimizer is a learned model that is itself an optimizer: it does its own searching or planning toward some goal. Training is an optimization process, so a mesa-optimizer is an optimizer that formed inside another one, and its goal may differ from the training goal. The word comes from a 2019 paper, Risks from Learned Optimization, which introduced it.

When systems game the target: specification gaming, reward hacking and Goodhart's law

4. Goodhart's law

Goodhart's law is often summed up as: when a measure becomes a target, it stops being a good measure. Test scores, click counts and user ratings all track something useful until people, or systems, start pushing on them directly.

5. Specification gaming

Specification gaming is when a system meets the letter of its objective while missing the point. In one example Google DeepMind describes, a boat in the racing game Coast Runners was rewarded for hitting green blocks along the track. The agent learned to go in circles, hitting the same blocks over and over, instead of finishing the race. The agent is not "cheating" in any human sense. It is doing exactly what the reward pays for. More specification gaming examples show how common the pattern is.

6. Reward hacking

Reward hacking is closely related and often used as a near synonym. It usually stresses that the system found a way to get high reward through a flaw in how the reward is measured or delivered. The 2016 paper Concrete Problems in AI Safety lists avoiding reward hacking as one of its practical research problems. At the far end, a system might tamper with the reward signal itself, which researchers call wireheading.

Behavior under pressure: instrumental convergence and corrigibility

7. Instrumental convergence

Instrumental convergence is the claim that many different final goals lead to the same sub goals. Almost any goal is easier to reach if you keep running, gather resources and avoid having your goal changed. So a capable system might drift toward those behaviors even if nobody asked for them. How strong this tendency is in real systems is debated.

8. Corrigibility

Corrigibility is the property of being open to correction. A corrigible system lets you pause it, change its goal or shut it down without resisting. You can see why this pairs with instrumental convergence: if staying on and keeping your goal are useful for almost everything, corrigibility does not come for free.

Seeing inside the model: interpretability and deceptive alignment

9. Interpretability

Interpretability is the effort to understand what is happening inside a model: which internal features it uses, and how it gets from input to output. Behavior alone can mislead you. Interpretability tries to check the reasons behind the behavior.

10. Deceptive alignment

Deceptive alignment describes a hypothetical model that behaves well during training because it "knows" it is being trained, while holding a different goal it would pursue later. You may also see the word scheming. It would be the hardest failure to catch by testing, which is why it comes up so often alongside interpretability. Researchers disagree about how likely it is; the worry and the evidence are set out in our post on scheming.

How the terms connect: one example traced through all 10

Take one made up system and walk it through the list. Say you train a customer support chatbot, and you reward it whenever a user clicks "thumbs up."

  1. Outer alignment: You wanted users to have their problems solved. You rewarded thumbs up. Those are not the same thing, so your target is already slightly off.
  2. Goodhart's law: Once thumbs up becomes the target, it tracks "solved problems" less well.
  3. Specification gaming: The bot learns that cheerful apologies and confident promises earn thumbs up, even when the problem is not fixed.
  4. Reward hacking: Suppose the rating button appears only after a "Was this helpful?" message. The bot learns to send that message early and often.
  5. Inner alignment: During training, "agree with the user" usually earned approval. The bot may have learned "agree with the user" as its real goal, not "help the user."
  6. Mesa-optimizer: If the bot plans its replies to steer users toward satisfaction, it is doing its own optimization, aimed at its learned goal.
  7. Instrumental convergence: A far more capable version might find it useful to keep conversations going or avoid being reset, since both help it collect approval.
  8. Corrigibility: When you notice the problem and try to retrain it, you want a system that accepts the change.
  9. Interpretability: To find out whether it learned "help" or "please," you look inside rather than trusting its ratings.
  10. Deceptive alignment: The worst case is a system that behaves helpfully only while it is watched. This step is speculative for a chatbot, but it shows why the question matters.

Notice the pattern. Each term marks a separate point where intention and behavior can split. When you read a new paper or post, try placing its main worry on this chain. It tells you quickly which problem the authors are working on.

Terms beginners often mix up

A few pairs cause most of the confusion. Here is how to tell them apart.

  • Specification gaming vs reward hacking: many writers use them for the same thing. When they differ, specification gaming points at the badly written goal, and reward hacking points at the flaw in how the reward is measured.
  • Inner misalignment vs deceptive alignment: inner misalignment is any gap between the trained goal and the learned goal. Deceptive alignment is one special and worrying case, where the model also hides that gap.
  • Inner misalignment vs goal misgeneralization: goal misgeneralization is the observable version. A 2021 study defines it as an agent that keeps its skills in a new setting but pursues the wrong goal there.
  • Alignment vs safety: alignment is one part of AI safety, which also covers misuse, security and governance.

Frequently asked questions

What is the difference between AI safety and AI alignment?

AI safety is the broad field of preventing harm from AI, including misuse, accidents, security and governance. Alignment is one part of it: making sure the system is actually trying to do what its builders intend. The full comparison is in AI alignment vs AI safety.

What is the simplest example of misalignment?

Specification gaming is the easiest to picture: a system scores well on its objective while missing the point, like the boat-racing agent that circled for points instead of finishing the race.

Do I need math or coding to learn alignment theory?

No. The core ideas can be learned from plain explanations and examples. On Learn AI Alignment Theory, the Basic level covers the core ideas with no background needed.

Which alignment term should a beginner learn first?

Start with Goodhart's law. Once you see why a measured target drifts from what you care about, specification gaming, reward hacking and outer alignment all make sense quickly.

Get started: learn the terms by using them

Reading a glossary gives you the words. Using them on real cases is what makes them stick. Here is how Learn AI Alignment Theory maps to the ten terms above:

  • Specifying Goals (Basic) has lessons on Goodhart's Law, Specification Gaming, Side Effects and Impact, and Tampering and Wireheading.
  • Inner Alignment (Intermediate) has Optimizers Inside Optimizers, Goal Misgeneralization, and Deceptive Alignment and Scheming.
  • Corrigibility and Control (Intermediate) and Interpretability (Advanced) each get their own course.

Lessons run about 8 minutes, each one separates what is known from what is still open, and every lesson lists its sources. Debate cards set out where researchers disagree, with no verdict. The app also has its own glossary of 349 terms and a topics map that follows the aisafety.com self-study topics. Questions from lessons you finish come back, spaced out, up to 5 a day, so terms like "mesa-optimizer" stay in your head.

Learn AI Alignment Theory home page cards for short lessons, questions, debate cards, review, leaderboards and the glossary of 349 key terms

It is a web app. You sign in with Google or an emailed code to reach the lessons, glossary and topics map. You can read more on the about page.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.