Human Compatible by Stuart Russell: the main ideas
Human Compatible by Stuart Russell is a 2019 book arguing that the way AI is usually built, around fixed goals that machines chase as hard as they can, could become dangerous as machines get more capable. Its full title is Human Compatible: Artificial Intelligence and the Problem of Control. This guide covers the book's main argument, Russell's three principles for safer machines, a short exercise to try them yourself, and what reviewers made of it.
What Human Compatible by Stuart Russell argues
Stuart Russell is a British computer scientist and a professor of computer science at the University of California, Berkeley. With Peter Norvig he wrote Artificial Intelligence: A Modern Approach, a textbook used in more than 1,500 universities in 135 countries. In 2016 he founded the Center for Human-Compatible Artificial Intelligence at Berkeley. Human Compatible came out on 8 October 2019.
Wikipedia's article on Human Compatible sums up its case in two parts. First, the risk to humanity from advanced AI is a serious concern, even though nobody knows how fast AI will progress. Second, there is a way to build machines that stay on our side, and the book proposes one.
Russell argues that progress in AI is going to continue whatever critics say, because of economic pressure. Human-level AI could be worth many trillions of dollars, and you can already see the pull in products like self-driving cars and personal assistant software. He also answers common arguments for dismissing AI risk, and suggests they persist partly through tribalism: some AI researchers hear talk of risk as an attack on their field.
Why Russell says the standard model is dangerous
Russell calls the usual approach the standard model. People give a machine a goal, and success means getting better and better at reaching it. He calls this dangerously misguided.
The trouble is that a goal written by people rarely says everything they care about. Russell explains it with old stories: "This is essentially the old story of the genie in the lamp, or the sorcerer's apprentice, or King Midas: you get exactly what you ask for, not what you want." A system chasing a goal, he notes, will often push anything the goal leaves out to extreme values. If one of those things matters to us, the result can be very bad.
His sharpest example is a thought experiment about a geoengineering robot that "asphyxiates humanity to deacidify the oceans". It reaches its goal. It just was never told that people need air.
If a machine built this way became superintelligent, Russell argues, it would likely not fully reflect human values, and the result could be catastrophic. Because nobody knows how long it will take to build such a machine, or how long safety research will take, he says that research should start as soon as possible. You can see the same family of problems in the post on specification gaming.
Russell's three principles
Russell's answer is what he calls provably beneficial machines. Instead of a fixed objective, the machine's true objective stays uncertain. It becomes more certain only as it learns more about people and the world. That doubt, ideally, keeps it from badly misreading what people want, and gives it a reason to ask, check and cooperate.

He sums it up in three principles. They are meant for the people who build machines, not as rules typed into the machine itself.
- The machine's only objective is to maximize the realization of human preferences. Preferences here are broad: in Russell's words, they "cover everything you might care about, arbitrarily far into the future."
- The machine is initially uncertain about what those preferences are. It must give some chance, even a tiny one, to every preference a person could logically have.
- The ultimate source of information about human preferences is human behavior. Any choice a person makes between options counts as evidence.
For the third principle, Russell looks at inverse reinforcement learning, where a machine works out what someone is aiming for by watching what they do. The post on inverse reinforcement learning explains how that works.
Why doubt could keep the off switch working
Wikipedia's AI alignment article describes a related proposal. A machine that is unsure of its objective has a reason to let a person turn it off: it can take being switched off as evidence that stopping is what the person wants. But the article lists limits. The reason only holds if the person is acting sensibly. And there is a trade-off: a machine that is very unsure is not much use, while one that is too sure may not let itself be turned off. Learning from behavior also assumes people act close to perfectly, which is not true for hard tasks. The post on corrigibility covers this problem in depth.
A worked example: the three principles in one request
This is a thinking exercise, not a description of how any real system works. It takes about 15 minutes and makes the book's idea concrete.
- Pick a request. Write one sentence you might give a home helper machine, for example: "Have the house warm by 7 in the morning."
- Play the standard model. List what a machine chasing only that sentence could do: run the heating at full power all night, ignore the bill, ignore an open window.
- List what the sentence left out. Write three things you care about that it never mentions, such as cost, noise, or a sleeping baby who should not be too hot.
- Apply principle 2. Write one question a machine that is unsure of your wishes would ask before acting.
- Apply principle 3. Write one thing it could learn just by watching you: do you open windows at night, or wear a sweater indoors?
Compare the two lists. The gap between them is the problem Russell is pointing at, and the questions are his proposed way across it.
How the book was received
Several reviewers agreed with it. In The Guardian, Ian Sample called it "convincing" and "the most important book on AI this year". Richard Waters of the Financial Times praised its "bracing intellectual rigour", and Kirkus Reviews called it "a strong case for planning for the day when machines can outsmart us". James McConnachie of The Times was mixed: "Its technical parts are too difficult, and its philosophical ones too easy. But it is fascinating and significant."
Others disagreed on two points. In Nature, David Leslie, an Ethics Fellow at the Alan Turing Institute, wrote that Russell "fails to convince that we will ever see the arrival of a 'second intelligent species'". In a New York Times opinion essay, Melanie Mitchell argued that a truly intelligent robot would naturally be "tempered by the common sense, values and social judgment without which general intelligence cannot exist". Their view is close to the debate in the post on the orthogonality thesis, which asks whether intelligence and goals come as a package.
For the other major book in this debate, see the guide to Superintelligence by Nick Bostrom.
Explore the race as fiction in Contain ASI
Russell's point about economic pressure is the setting of Contain ASI. Its home page reads: "January 2024. Four AI labs. One race. For three years, keep every lab from crossing into recursive self-improvement before 2027." In a January 2025 Newsweek article, Russell himself wrote: "In other words, the AGI race is a race towards the edge of a cliff."

It is a story strategy game for 1 to 4 players that runs in your browser, where you play as a researcher, research manager, CEO or government. It is fiction, not a forecast: the footer says it is inspired by the AI 2027 scenario and that all its labs, people and events are made up.
Frequently asked questions
What is the main idea of Human Compatible?
That building AI to chase fixed goals is dangerous, and that machines should instead be unsure of what people want and learn it from how people behave.
What are Stuart Russell's three principles?
The machine's only objective is to realize human preferences; it starts out uncertain what those preferences are; and it learns them from human behavior.
Is Human Compatible hard to read?
Reviewers called it accessible and witty, though The Times found its technical parts too difficult and its philosophical ones too easy.
Is Contain ASI based on the book?
No. The game's page names the AI 2027 scenario as its inspiration and says all its labs, people and events are fictional.
Get started
Want to feel the pressure Russell describes from the inside? Play Contain ASI: a story strategy game for 1 to 4 players that runs in your browser. Its labs, people and events are all fictional. You will need a recent Chrome, Edge, Firefox or Safari, because it runs on WebGL 2.
Comments
No comments yet.
Sign in or make an account to comment.