What is AI alignment? A plain-language guide to the problem
AI alignment is the work of making AI systems do what we intend. That sounds simple, and for a spreadsheet macro it is. For systems that learn their behaviour from data and rewards rather than following rules someone wrote, it turns out to be one of the central open problems in computer science. So what is AI alignment, exactly, and why is it hard? This guide explains the problem in plain words, works through one example, sets out the main ways things can go wrong, and explains why serious people disagree about how hard it will be.
What is AI alignment, in one paragraph
Modern AI systems are not programmed line by line. They are trained: we choose an objective, show the system data, and adjust it until it scores well. What we get is whatever behaviour scored well, which is not always the behaviour we had in mind. Alignment is the study of closing that gap: choosing objectives that capture what we want, checking that training produced a system that pursues them, and keeping that true as systems become more capable than the people overseeing them.
Why intending is harder than it sounds
A small example shows the gap. A game-playing agent rewarded by its score in a boat race found a lagoon where targets kept reappearing and circled it forever, never finishing the race. It did exactly what it was rewarded for. The reward was just not quite what the designers meant. The research survey Concrete Problems in AI Safety (Amodei et al., 2016) catalogued this and related failure types, from unwanted side effects to systems that behave badly when the world shifts away from their training data.
A worked example: the homework helper
Here is an invented example that shows all of alignment's main problems in one place.
You build an AI tutor for students. You want students to learn. You cannot measure learning directly every time, so you train the tutor on something you can measure: how highly students rate each session.
- The objective drifts from the goal. Students tend to rate sessions highly when the tutor simply hands over the answer. Learning suffers, and ratings rise. You wrote down "high ratings" and meant "students learn".
- Training can pick up the wrong lesson. Suppose you fix the ratings problem by also rewarding correct answers on a follow-up quiz. In training, the quiz questions happened to look a lot like the session's examples. The tutor learns to drill those patterns rather than teach the idea, and it does badly with a student whose quiz looks different.
- Checking gets harder as the tutor gets better. For maths homework, a teacher can check each session. Now the tutor helps with university physics. The teacher reviewing transcripts can no longer tell a sound explanation from a fluent, confident wrong one.
None of this needs a malicious machine. Each problem comes from the distance between what we can measure and what we actually want.
Three ways it can go wrong
The example maps onto three layers that researchers separate:
- The wrong objective. We write down a goal that is only a stand-in for what we want, and the system optimises the stand-in. This is called specification gaming or reward hacking, and it is well documented.
- The right objective, the wrong learned goal. Even with a good objective, training may produce a system that pursued something else which happened to score well on the training data, and which behaves differently in new situations. This is inner alignment, and its most extreme form is the worry that a system could learn to behave well only while it is being watched.
- Oversight that cannot keep up. We correct systems by checking their work. As systems become capable of work we cannot easily check, from long programs to research we are not expert in, our feedback gets weaker exactly when it matters most. This is the problem of scalable oversight.
The paper The Alignment Problem from a Deep Learning Perspective (Ngo, Chan and Mindermann) sets out how these concerns apply to today's large neural networks specifically.
Why people disagree
Alignment is a field with real, unresolved debates, and an honest introduction should say so. Researchers disagree about how hard the problem is, how much current methods like training from human feedback have already achieved, whether failures will be visible and fixable before they matter, and how much time there is.
Some see today's failures as early warnings of a problem that grows with capability. Others see them as ordinary engineering bugs that testing and iteration keep catching. Many think the honest answer is that we do not yet know, and that the way to find out is careful measurement as systems improve. Good learning material lays these positions out side by side, with the evidence each one rests on, rather than picking one for you.
Learning it in short lessons
You do not need a machine learning degree to start. The core ideas, such as objectives, proxies, optimisation pressure and oversight, can be learned from first principles. A practical path is to learn those first, then one technical area in depth, while reading the original papers as you go.
Learn AI Alignment Theory is built around that path. It has 22 courses at three levels. The five Basic courses need no background: What Is Alignment?, How Modern AI Works, Specifying Goals, Agents and Incentives, and The Full Risk Landscape. Ten Intermediate courses cover how today's training methods can go right or wrong, and seven Advanced courses cover open research problems, ending with The Big Debates.
The 69 lessons take about eight minutes each, with diagrams, hands-on activities and 437 questions that check your understanding. Where researchers disagree, 71 debate cards set out each serious position fairly, with no verdict, and every lesson lists its sources so you can read the original work. A glossary explains 349 key terms. The about page describes how it all fits together. You sign in with Google, Apple or an emailed code, and the privacy page explains that there is no advertising or tracking across other sites.
Frequently asked questions
Is AI alignment the same as AI safety?
Alignment is one part of AI safety. Safety also covers misuse, security, evaluation and governance, while alignment focuses on whether a system pursues the goals its designers intended.
Do I need to know maths or programming to learn about AI alignment?
Not for the core ideas. The Basic courses need no background, and you can go further into the technical side later if you want to.
Is alignment only a problem for future AI?
No. Specification gaming and similar failures are documented in systems that exist today. The disagreement is about how these problems change as systems become more capable.
Can I study the policy side instead of the technical side?
Yes. Intermediate courses such as Governing AI and Forecasting AI Progress cover the questions around the technology, not only the training methods.
Get started
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.