Eliezer Yudkowsky: who he is and what he argues about AI
Eliezer Yudkowsky is an American artificial intelligence researcher and writer on decision theory and ethics. He is known for popularizing the idea of friendly AI, and for his warning, in the subtitle of his 2025 book, of "Why Superhuman AI Would Kill Us All". This post explains who he is, what he argues, the main objection to one of his key ideas, and where to start reading. The claims come from Wikipedia's Eliezer Yudkowsky article, with one date from its article on the institute he founded.
Who Eliezer Yudkowsky is
Yudkowsky was born in 1979. In 2000 he founded the Singularity Institute for Artificial Intelligence, which in 2013 took the name Machine Intelligence Research Institute (MIRI). MIRI is a private research nonprofit based in Berkeley, California, and he is a research fellow there.
His influence reaches beyond his own institute. His work on a runaway intelligence explosion influenced the philosopher Nick Bostrom's 2014 book Superintelligence; see our post on Bostrom's Superintelligence. His views are also discussed in Stuart Russell and Peter Norvig's undergraduate textbook, Artificial Intelligence: A Modern Approach.

Friendly AI: build it safe from the start
The idea he is known for popularizing is friendly AI: an AI with a desire not to harm humans. The textbook cites his 2008 argument that friendliness should be designed in from the start. At the same time, the designers should accept two things: their own designs may be flawed, and the system will learn and change over time.
So, in his framing, the real task is mechanism design: building a way for an evolving AI to stay under checks and balances, with goals that remain friendly as it changes. For the history of the term, see our post on friendly AI.
Safe defaults when goals go wrong
He also responds to a worry called instrumental convergence: an AI with badly designed goals would, by default, have reasons to mistreat humans. Our post on instrumental convergence explains it. Yudkowsky and other MIRI researchers recommend work on AI agents that settle on safe default behaviour even when their goals are specified wrongly.
Coherent extrapolated volition
In 2004 he proposed coherent extrapolated volition: AIs designed to pursue what people would want under ideal conditions of knowledge and morality, not what they happen to want today.
The intelligence explosion, and the "village idiot and Einstein" point
The intelligence explosion is a scenario first set out by I. J. Good. Recursively self-improving AI systems would quickly go from below human general intelligence to superintelligence. Our post on the intelligence explosion covers the idea and the disagreement around it.
Bostrom's book cites Yudkowsky on why people may misread such a jump. In Yudkowsky's words: "AI might make an apparently sharp jump in intelligence purely as the result of anthropomorphism, the human tendency to think of 'village idiot' and 'Einstein' as the extreme ends of the intelligence scale, instead of nearly indistinguishable points on the scale of minds-in-general."
In plain words: to us, the gap between the least and most clever person feels huge. On a scale that includes all possible minds, it may be tiny. So an AI could pass from "village idiot" to "Einstein" in what looks like a sudden leap.
The main objection
Russell and Norvig raise a counterpoint in their textbook. Computational complexity theory, the study of how hard problems are to solve, shows known limits to how efficiently algorithms can solve various tasks. If those limits are strong, an intelligence explosion may not be possible at all.
Calling for a halt
In a 2023 op-ed for Time magazine, Yudkowsky argued for international agreements to limit AI, including an indefinite worldwide moratorium on large AI training runs. He went further: he suggested that participating countries should be willing to take military action to enforce it, such as "destroy[ing] a rogue datacenter by airstrike".
According to the article, the op-ed helped bring the debate about AI alignment into the mainstream. It led a reporter to ask the US president a question about AI safety at a press briefing.
If Anyone Builds It, Everyone Dies
With Nate Soares, Yudkowsky wrote If Anyone Builds It, Everyone Dies, published on 16 September 2025 and a New York Times bestseller. Our post on If Anyone Builds It, Everyone Dies sets out its argument and its critics.
Rationality writing
Alongside AI, Yudkowsky writes about how to think well.
- Overcoming Bias: between 2006 and 2009, he and Robin Hanson were the main writers on this blog, sponsored by Oxford University's Future of Humanity Institute.
- LessWrong: in February 2009 he founded it as a "community blog devoted to refining the art of human rationality". It has grown into the main online hub for the rationalist community.
- The Sequences: over 300 of his blog posts were released by MIRI in 2015 as the ebook Rationality: From AI to Zombies.
- Inadequate Equilibria: a 2017 ebook on why societies stay inefficient.
- Fiction: Harry Potter and the Methods of Rationality, a fan novel that uses the Harry Potter series to illustrate science and rationality.
A worked example: a reading path through his ideas
If you want to understand Yudkowsky's view without starting with a 300-post ebook, follow this path over four sittings. Hold one question for each step.
- Friendly AI. Read our friendly AI post. Question: why must safety be designed in from the start, rather than added later?
- Instrumental convergence. Read the instrumental convergence post. Question: why would an AI with badly designed goals have reasons to mistreat humans by default?
- The intelligence explosion. Read the intelligence explosion post, then reread the "village idiot and Einstein" quote above. Question: is the human range of intelligence wide or narrow? Then weigh Russell and Norvig's objection.
- The call for a halt. Read the If Anyone Builds It post. Question: if steps 1 to 3 hold, what follows, and what do critics say?
By the end, you can state his case in four sentences and name the main objection. That is enough to follow almost any debate where his name comes up.
Exploring the idea in Contain ASI
Contain ASI is a story strategy game for 1 to 4 players that runs in your browser and needs WebGL 2. Chapter 1, Before ASI, starts in January 2024 with four fictional AI labs in one race. As a researcher, research manager, CEO or government, you try for three years to keep every lab from crossing into recursive self-improvement before 2027.
It is a game, not a forecast: the page says all its labs, people and events are fictional, and the page does not name Yudkowsky. But it puts you in the position his call for a halt describes: trying to hold a race back before anyone builds superintelligence.
Frequently asked questions
Who is Eliezer Yudkowsky?
He is an American AI researcher and writer on decision theory and ethics, known for popularizing friendly AI. He founded the institute now called MIRI and is a research fellow there.
What is Eliezer Yudkowsky's main argument about AI?
His 2025 book's subtitle sums it up: "Why Superhuman AI Would Kill Us All". He argues friendliness must be designed in from the start, and in 2023 he called for an indefinite worldwide moratorium on large AI training runs.
What did Yudkowsky write?
His works include If Anyone Builds It, Everyone Dies with Nate Soares, Rationality: From AI to Zombies, Inadequate Equilibria and the fan novel Harry Potter and the Methods of Rationality.
Did Yudkowsky found LessWrong?
Yes. He founded LessWrong in February 2009, and it grew into the main online hub for the rationalist community.
Get started
Want to see how hard it is to stop a race? Open Contain ASI in your browser and start a campaign as a researcher, research manager, CEO or government.
Comments
No comments yet.
Sign in or make an account to comment.