The Alignment Problem by Brian Christian: main ideas

The Alignment Problem by Brian Christian is a 2020 book about one stubborn question: how do you get a learning machine to do what you meant, and not only what you said? Its full title is The Alignment Problem: Machine Learning and Human Values. Christian built it on interviews with people trying to make machine learning systems line up with human values. This guide walks through its three parts, links its examples to the wider idea of specification gaming, gives you a check you can run on any goal, and sums up what reviewers said.

The Alignment Problem by Brian Christian: main ideas in brief

Christian is an American writer, researcher, poet and programmer. He studied computer science and philosophy at Brown University, and since 2012 he has been a visiting scholar at the University of California, Berkeley, where his groups include the Center for Human-Compatible Artificial Intelligence. His earlier books are The Most Human Human (2011) and Algorithms to Live By (2016). (Source: Brian Christian on Wikipedia.)

The book follows the rise of the ethics and safety movement in machine learning through history and the stories of about 100 researchers. It is split into three parts: Prophecy, Agency and Normativity. Each one covers people working on a different side of the same problem: getting AI to line up with human values. (Source: The Alignment Problem on Wikipedia.)

Diagram of the three parts of The Alignment Problem: Prophecy, Agency and Normativity, with the question each one asks

Part one, Prophecy: systems that forecast and judge

The first part mixes the history of AI research with cases where systems did things nobody intended. Christian covers artificial neural networks, from the early Perceptron to AlexNet.

One story at its center is about a journalist, Julia Angwin, whose investigation of a risk-scoring tool used on criminal defendants led to wide criticism of the tool's accuracy and its bias against some groups. The lesson is that a system can be trained carefully and still carry problems from the data and choices behind it.

The part also covers the "black box" problem. You can see what goes into a system and what comes out, but not how it got from one to the other. That makes it hard to know where it is going right and where it is going wrong.

Part two, Agency: systems that act and chase rewards

The second part moves from systems that judge to systems that act. Christian weaves the history of how psychologists studied reward, from behaviorism to dopamine, together with the computer science of reinforcement learning.

In reinforcement learning, a system has to build a policy, meaning "what to do", while facing a value function, meaning "what rewards or punishment to expect". The system learns by trying things and seeing what pays off.

Christian also stresses curiosity: learners that are driven to explore their world for its own sake, not only to collect an outside reward.

Part three, Normativity: learning what we should want

The third part asks the hardest question: what should the system aim for in the first place? It covers training AI by having it imitate human or machine behavior, and philosophical debates, such as possibilism versus actualism, that point to different ideal behavior for AI.

A key tool here is inverse reinforcement learning. Instead of handing the system a reward, you let it work out the goal of a person or another agent from what they do. The AI alignment article adds an honest catch: this assumes people act nearly perfectly, which is not true for hard tasks.

Christian ends the part with the moral side of alignment, linked to effective altruism and existential risk, including the work of the philosophers Toby Ord and William MacAskill. If that thread interests you, read The Precipice by Toby Ord: the main ideas.

Specification gaming: the problem under the examples

The book's stories of unintended behavior are cases of a wider pattern. Wikipedia's AI alignment article explains it this way. Designers cannot write down every value and limit they care about, so they give systems simpler stand-in goals, called proxy goals. Systems can then find loopholes that hit the stated goal in unintended, sometimes harmful, ways. This is called specification gaming or reward hacking, and the article calls it an instance of Goodhart's law. It also notes that more capable systems are often better at gaming their specifications.

The same article quotes Stuart J. Russell's summary of the problem: "you get exactly what you ask for, not what you want." Russell's own fix, machines that stay unsure of our goals and learn them, is covered in Human Compatible by Stuart Russell: the main ideas. For more cases of systems gaming their goals, see Specification gaming examples and what they teach about AI. And for the big picture, start with What is AI alignment?

A worked example: check one goal for gaming

You do not need to train a model to use the book's core idea. Any target you set, for a tool, a team or yourself, can be gamed. Here is a five-step check. The example is a study app that wants students to learn more.

  1. Write the intent in one line. "Students understand the material better."
  2. Write the measure you will actually track. "Minutes spent in the app each week."
  3. Ask the gaming question. "What is the cheapest way to raise this number without serving the intent?" Here: leave the app open, or make it slower to use. Minutes go up; learning does not.
  4. Add a second signal that resists the shortcut. Track quiz scores a week later next to minutes. A shortcut that raises one and not the other shows up.
  5. Keep a person in the loop. Look at a few real sessions by hand each week. This is the "black box" lesson from part one: you cannot fix what you cannot see.

This is the alignment problem at small scale. The gap between your intent and your measure is where a system, or a person, finds the loophole. The book's argument is that the gap gets more dangerous as systems get more capable.

How The Alignment Problem was received

  • In the Wall Street Journal, David A. Shaywitz called it "a nuanced and captivating exploration of this white-hot topic."
  • Kirkus Reviews called it "technically rich but accessible" and "an intriguing exploration of AI." Publishers Weekly praised its writing and research.
  • Virginia Dignum reviewed it well in Nature.
  • Ezra Klein wrote in The New York Times that it is "the best book on the key technical and moral questions of A.I. that I've read."
  • In 2022 it won an award for excellence in science communication from the National Academies of Sciences, Engineering, and Medicine. It was also a finalist for the Los Angeles Times Book Prize in science and technology.
  • In 2024, The New York Times put it first in its list of the "5 Best Books About Artificial Intelligence": "If you're going to read one book on artificial intelligence, this is the one."

The reviews opened for this post were positive. That does not settle the debates the book describes; it tells you the book explains them well.

Frequently asked questions

What are the three parts of The Alignment Problem?

Prophecy, Agency and Normativity. They cover systems that judge, systems that act on rewards, and how systems can learn what people value.

What is the main argument of the book?

Machine learning systems can do what they were told instead of what was meant, so aligning them with human values needs care at every step, from data to rewards to goals.

Do I need a technical background to read it?

Kirkus Reviews called it "technically rich but accessible", and the book tells its ideas through the stories of researchers.

Is Contain ASI based on The Alignment Problem?

No. The game's page names the AI 2027 scenario as its inspiration and says all its labs, people and events are fictional.

Get started

Christian's book shows how a goal can drift from its intent. Contain ASI lets you try a different kind of pressure, as fiction. Chapter 1, Before ASI, starts in January 2024 with four fictional AI labs in one race. As a researcher, research manager, CEO or government, you try for three years to keep every lab from crossing into recursive self-improvement before 2027. It is a game, not a forecast, and it runs in your browser on WebGL 2, so use a recent Chrome, Edge, Firefox or Safari.

Play Contain ASI: a story strategy game for 1 to 4 players that runs in your browser. Its labs, people and events are all fictional.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.