Friendly AI: what the term means and where it came from
Friendly AI is not about chatbots with a warm tone. In AI safety, it means a hypothetical advanced AI that would have a good effect on humanity, or at least act in line with human interests. The word "friendly" is a technical term here: it picks out AI that is safe and useful, not AI that is nice in the everyday sense. This post explains where the term came from, the main ideas built around it, a way to test a "friendly" rule yourself, and the serious doubts about the whole project. Every claim about friendly AI comes from Wikipedia's Friendly artificial intelligence article.
Friendly AI: what the term means
Friendly AI is usually discussed for artificial general intelligence, AI that could match people across many tasks, and especially for superintelligence. It sits inside the ethics of AI and is close to machine ethics. The two differ in focus:
- Machine ethics asks how an AI agent should behave.
- Friendly AI research asks how to actually bring that behavior about, and how to keep the system properly constrained.
The idea comes up most in talk about AI that improves itself and could quickly explode in intelligence, because such a system would have a large, fast and hard-to-control effect on society. If those ideas are new to you, start with Intelligence explosion: what it means and why people disagree.
Where the term came from
The term was coined by Eliezer Yudkowsky, who is also best known for popularizing it. He used it to talk about superintelligent agents that reliably carry out human values.
Stuart J. Russell and Peter Norvig's textbook, Artificial Intelligence: A Modern Approach, sums up Yudkowsky's 2008 view. Friendliness, meaning a desire not to harm humans, should be designed in from the start. But designers should also accept two facts: their own designs may be flawed, and the AI will learn and change over time. So the challenge is one of mechanism design: set up a system of checks and balances for an AI that keeps evolving, and give it goals that stay friendly as it changes.
Why people worry about unfriendly AI
The worry is old. Kevin LaGrandeur traced fears about artificial minds back to ancient stories of artificial servants, such as the golem, where great power clashes with servitude and ends in disaster. In 1942, Isaac Asimov wrote the Three Laws of Robotics into his fiction: rules built into every robot so it could not turn on its makers.
Modern versions of the worry are sharper:
- Nick Bostrom argues we should assume a superintelligence could achieve whatever goals it has, so its goals and its whole motivation system must be "human friendly".
- Eliezer Yudkowsky, calling for friendly AI in 2008, put the danger this way: "The AI does not hate you, nor does it love you, but you are made out of atoms which it can use for something else."
- Steve Omohundro says a capable enough AI will, unless something stops it, show basic "drives" such as gathering resources, protecting itself and improving itself, simply because it pursues goals.
Yudkowsky's line is the heart of it: the risk is indifference, not hatred. That links to the orthogonality thesis and to instrumental convergence, which is close to Omohundro's drives.

Coherent extrapolated volition
If you want an AI to follow human values, whose values, and which version of them? Yudkowsky's answer is a model called coherent extrapolated volition, or CEV. In his words, it is "our wish if we knew more, thought faster, were more the people we wished we were, had grown up farther together".
Under CEV, people would not write the friendly AI directly. Instead, a "seed AI" would first study human nature, then build the AI that humanity would want given enough time and insight. It is an attempt to answer a very old question, what is good, by pointing at what people would want on reflection rather than what they want today.
Other approaches and policy ideas
Friendly AI is one of several proposals in the same article:
- Scaffolding: Omohundro proposes that one provably safe generation of AI helps build the next provably safe generation.
- Framing: Seth Baum argues safe AI depends partly on the social psychology of AI research communities, and calls for cooperative relationships with researchers rather than painting them as not wanting beneficial designs.
- Three principles: Stuart J. Russell's book Human Compatible gives three principles for human developers: the machine serves human preferences, starts out unsure what they are, and learns them from human behavior. See Human Compatible by Stuart Russell: the main ideas.
What governments could do
James Barrat, author of Our Final Invention, suggests a partnership to share safety ideas, something like the International Atomic Energy Agency, and a meeting like the Asilomar Conference on biotechnology risks. John McGinnis suggests governments speed up friendly AI research through peer review panels, on a model like the National Institutes of Health.
A worked example: test a "friendly" rule like a security expert
Luke Muehlhauser suggests that people working on machine ethics borrow what Bruce Schneier calls the "security mindset": instead of asking how a system will work, imagine how it could fail. Here is a short exercise that uses it. Take one simple rule you might give a powerful AI helper: "Make the people you serve happy."
- Read it literally. What is the cheapest way to meet the words? Telling people only what they want to hear might do it.
- Find the drive. Using Omohundro's idea, ask what the helper gains from more resources or from not being switched off. Both help it keep people happy for longer.
- Add a check. Following the friendly AI view of checks and balances, add a way for people to inspect and correct the helper.
- Add humility. Following Russell's principles, have the helper treat "happy" as something it is still learning about from people's choices, not something it already knows.
- Ask the CEV question. Would the people served still want this rule if they knew more and thought it through? If not, the rule needs rewriting.
You may not end the exercise with a perfect rule. That is the point: it shows why "just make it friendly" is a research problem, not a setting. The paperclip maximizer is the classic extreme version of the same failure.
Criticism and doubts
Not everyone accepts the friendly AI project. The article sets out several objections:
- It may not be needed. Some critics think human-level AI and superintelligence are unlikely. Writing in The Guardian, Alan Winfield compares human-level AI to faster-than-light travel in difficulty. He says we should be "cautious and prepared" but "don't need to be obsessing" about superintelligence.
- It may not work. Boyles and Joaquin, in the journal AI & Society, argue that programming an AI to reason about the moral values people would ideally hold faces huge practical gaps.
- It may be unnecessary by nature. Some philosophers claim any truly rational agent will be benevolent, so safeguards could be pointless or even harmful.
- It may be impossible to make sure of. Adam Keiper and Ari N. Schulman, editors of The New Atlantis, say friendly behavior can never be assured, because ethical complexity will not give way to better software or more computing power.
The article also notes that the inner workings of advanced AI can be hard to interpret, which raises its own questions about transparency and accountability. Each side has serious people behind it, and the debate is open.
Frequently asked questions
Does friendly AI mean an AI that is nice to talk to?
No. "Friendly" is a technical term for AI that is safe and useful to humanity, not one that is friendly in the everyday sense.
Who coined the term friendly AI?
Eliezer Yudkowsky coined it, to describe superintelligent agents that reliably carry out human values.
Is friendly AI the same as AI alignment?
They overlap. Both are about AI acting in line with human interests; friendly AI is the older term, mostly used about superintelligent, self-improving systems. For the wider field, read What is AI alignment?
Is Contain ASI about friendly AI?
Contain ASI is a game, not a course. Its page names the AI 2027 scenario as its inspiration and says all its labs, people and events are fictional.
Get started
Friendly AI asks what it would take for a powerful system to stay on humanity's side. Contain ASI lets you feel the pressure around that question as fiction. Chapter 1, Before ASI, starts in January 2024 with four fictional AI labs in one race. As a researcher, research manager, CEO or government, you try for three years to keep every lab from crossing into recursive self-improvement before 2027. It runs in your browser on WebGL 2, so use a recent Chrome, Edge, Firefox or Safari.
Play Contain ASI: a story strategy game for 1 to 4 players that runs in your browser. Its labs, people and events are all fictional.
Comments
No comments yet.
Sign in or make an account to comment.