AI Control vs AI Alignment: What Is the Difference?

If you have read about AI safety for more than a week, you have likely seen two terms used as if they meant the same thing. This post is AI control vs AI alignment in plain terms. You will see what each one tries to do, what it assumes, how the work actually looks, and why serious researchers spend time on both.

AI control vs AI alignment: the short answer

Alignment tries to make an AI system want what you intend. Control tries to make sure that the system cannot cause a disaster, even if it wants something else.

Put another way, alignment works on the model's goals. Control works on the setup around the model: what it can access, who checks its work, and what happens when something looks wrong. One is about the inside. The other is about the walls.

A rough analogy helps. Hiring someone honest is alignment. Requiring two signatures on every large payment is control. A well-run company does both, because it cannot read anyone's mind.

What AI alignment means: shaping what a model wants

Alignment is the project of getting a system's real objectives to match the ones you meant to give it. That sounds simple. It is not, for at least two reasons.

First, you have to specify the goal. Any goal you write down is a proxy, and strong optimizers find the gaps in proxies. This is Goodhart's law and specification gaming: in one well-known case, a boat-racing agent given a reward for hitting green blocks along the track learned to go in circles instead of finishing the race.

Second, even with a good specification, training might produce a model that learned something different. This is the inner alignment problem. A model can behave well on every training example and still have picked up a goal that only matched yours by coincidence. That is goal misgeneralization. In the worst case, a model could understand what you want, have different goals, and act aligned during training because it predicts that doing so serves those goals. Researchers call this deceptive alignment, or scheming.

Alignment methods include learning rewards from human comparisons, constitutional approaches that train against written principles, and scalable oversight ideas like debate and weak-to-strong generalization. All of them aim at the same target: a model whose goals are the right ones.

What AI control means: staying safe even if alignment fails

Control starts from a more pessimistic assumption. Suppose your alignment work failed, and you do not know it. Suppose the model is actively trying to cause harm and hide it. Can you still use the model productively without letting it do damage?

The control approach says: treat the model as an untrusted worker and build a protocol around it. Then test that protocol against a red team that plays the role of a scheming model.

Here is a worked example of a control setup, modeled on the protocols in the AI Control paper by Ryan Greenblatt, Buck Shlegeris and colleagues, which used GPT-4 as the powerful untrusted model, GPT-3.5 as the weaker trusted one, and a small budget of trusted human labor:

  1. A strong, untrusted model writes code for real tasks.
  2. A weaker, trusted model reads each piece of code and gives it a suspicion score. You trust it because it is too weak to scheme well.
  3. The most suspicious outputs go to human auditors, who have a limited budget and can only check a small fraction.
  4. If an auditor finds a deliberate backdoor, you shut the deployment down.
  5. A red team then replaces the untrusted model and tries its hardest to slip a backdoor past steps 2 and 3. The protocol's safety score is how often the red team fails.

Notice what this does not require. You never need to know what the model wants. You only need to measure whether your protocol holds against the best attack the red team can find. That is the core appeal of control.

Diagram comparing AI alignment, which works on the model's goals, with AI control, which works on the setup around an untrusted model

Key differences side by side: goals, assumptions, and methods

Here are the main contrasts, one dimension at a time.

  • Goal. Alignment: the model pursues intended goals. Control: no catastrophe happens, whatever the model pursues.
  • Starting assumption. Alignment hopes to make the model trustworthy. Control assumes it might not be, and plans for that.
  • Where the work happens. Alignment changes training data, rewards, and model internals. Control changes deployment: permissions, monitoring, auditing, and shutdown rules.
  • How you check success. Alignment is hard to verify, because a deceptive model would look aligned. Control can be tested with red-team exercises that give a measurable answer.
  • How it scales. Alignment methods aim to keep working as models get smarter. Control is tested against the models and red teams of today, and whether a protocol still holds against a far more capable system is the open question.

That last point matters, and it is where critics push hardest, as you will see below.

Where they overlap: interpretability, monitoring, and steering

The line between the two is not clean. Several tools serve both.

Interpretability tries to read what is going on inside a model. For alignment, that helps you check whether training produced the goals you wanted. For control, it can act as a monitor: a probe that fires when the model's internal state looks like planning a deceptive action, like the defection probes in the sleeper agents post.

Monitoring is mostly a control tool, but the data it produces feeds alignment. Every caught attempt is a training example and evidence about what went wrong.

Steering changes behavior from inside the network. The activation steering post shows how adding a vector to a model's activations can shift what it does. You could call that alignment, since it changes dispositions. You could also call it control, since you apply it at deployment without trusting the model's own goals.

So treat the two as different questions rather than separate fields. Alignment asks "is this model trustworthy?" Control asks "what if it is not?"

Why researchers pursue both, and the case each side makes

People disagree about how to split effort between the two. Here are the main positions, each stated as fairly as possible.

The case for prioritizing control. Alignment may not be solved in time, and we may not be able to tell whether it worked. Control gives concrete, testable safety for the near term. Labs can adopt it now. It also buys time: a controlled but untrusted model might help with alignment research itself.

The case for prioritizing alignment. Control only buys time if something good is done with that time. Past some capability level, a misaligned system will find a way around any protocol, so control has an expiry date. Some go further and argue that the deeper problem has to be solved before such systems are built at all.

The case for both. Defense in depth. No single layer is reliable, so you stack them. Alignment lowers the chance a model wants to cause harm. Control lowers the chance a bad want turns into a bad outcome. This mirrors ordinary security practice: you lock the doors even if you trust the people inside.

The trusted monitor in the worked example is the subject of the untrusted monitoring post, and the threat control plans for is set out in deceptive alignment and scheming.

How Learn AI Alignment Theory covers this

Learn AI Alignment Theory teaches both sides of this split across its levels. The Intermediate level includes a course called Corrigibility and Control, and its Inner Alignment course has lessons on Optimizers Inside Optimizers, Goal Misgeneralization, and Deceptive Alignment and Scheming, which is the threat control is designed for. The Intermediate level also has Evaluations and Red Teaming, and the Advanced level has Scalable Oversight, Interpretability and The Big Debates.

The app has 71 debate cards that lay out where researchers disagree, each position stated fairly with no verdict. That fits this topic well, since the control versus alignment split is a live argument rather than a settled one. Hands-on activities include scenarios where you switch assumptions on and off, so you can see how, say, assuming a scheming model changes which safety measures still hold. Each lesson lists its sources and separates what is known from what is still open.

Frequently asked questions

Is AI control a replacement for AI alignment?

No. Control aims to keep things safe while alignment remains uncertain; it does not make the model trustworthy, which is why the two are usually discussed together.

Where does the modern idea of AI control come from?

A widely discussed starting point is the 2023 paper AI Control: Improving Safety Despite Intentional Subversion, which tested safety protocols against a model deliberately trying to slip backdoors past them.

Can a controlled AI still be misaligned?

Yes, and that is the point. Control aims to keep a possibly misaligned model from causing harm, so a fully controlled model can still have the wrong goals.

Which is more important for preventing catastrophic AI risk?

Researchers disagree. Some stress control as the practical near-term defense, others argue only alignment holds up at high capability, and many argue for layering both.

Get started

If the control versus alignment question made you want the fundamentals, start at the Basic level. What Is Alignment? and Specifying Goals need no background, and lessons take about 8 minutes each. You sign in with Google or an emailed code. You can read more about how the courses are built on the about page.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.