Asimov's Three Laws of Robotics and AI safety
The Three Laws of Robotics are the most famous answer to the question "how do we keep machines from hurting us?" The science fiction author Isaac Asimov wrote them in 1942. They are short, clear and easy to remember. They are also, as Asimov's own stories show again and again, full of holes. This post explains the laws, the Zeroth Law he added later, the loopholes, and why people who work on AI safety today doubt that rules like these can work. Every claim comes from Wikipedia's Three Laws of Robotics article.
The Three Laws of Robotics, word for word
Asimov introduced the laws in his 1942 short story "Runaround", later collected in I, Robot (1950). In the stories they are quoted from a fictional "Handbook of Robotics, 56th Edition, 2058 A.D.":
- A robot may not injure a human being or, through inaction, allow a human being to come to harm.
- A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.
- A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.
The order matters. Each law gives way to the ones above it. In Asimov's fiction, the laws are built into almost all of his robots as a safety feature and cannot be bypassed.
Why Asimov wrote them
When Asimov started writing in 1940, he felt that "one of the stock plots of science fiction was ... robots were created and destroyed their creator." He wanted something else. He decided his robots would not "turn stupidly on his creator" just to repeat the old warning tale.
So the laws were a storytelling choice: a way to write robots that were safe by design, and then find the drama in the edge cases. Many of his robot stories turn on robots behaving in strange, counter-intuitive ways as an unintended result of how they apply the laws. Wikipedia notes that the laws have since influenced thinking on the ethics of artificial intelligence.
The Zeroth Law: protecting humanity
In later books, where robots help run whole planets, Asimov added a fourth law and put it first. It is called the Zeroth Law, because lower numbers come first:
A robot may not injure humanity or, through inaction, allow humanity to come to harm.
This sounds like an upgrade. In the stories it is also a danger. The robot R. Giskard Reventlov, acting on it, comes to accept harming individual people in the service of the abstract idea of humanity. A rule meant to make a robot safer gives it a reason to override the rule that protected each person. That is close to the worry in value lock-in: a powerful system acting on its own idea of what is good for everyone.

The loopholes Asimov found himself
Asimov's stories are, in a sense, a long list of ways the laws fail:
- Not knowing. In The Naked Sun, a detective points out that robots can break any law without knowing it. A robot could be told to add something to a person's food without knowing it is poison. A clever criminal could split a harmful task among many robots so none of them sees the harm.
- Who counts as human. The laws assume "human being" is well defined. In one society, robots are told that only people with a certain accent are human, so they can harm everyone else with no conflict.
- Dilemmas. In the story "Liar!", a robot will hurt people if it tells them something and hurt them if it does not. Robots forced into a spot where they cannot obey the First Law can suffer an irreversible mental collapse.
- Weighing harms. More advanced robots weigh their options and break the laws as little as possible rather than do nothing. That makes them useful, as robot surgeons for example, but it also means "do no harm" has become "do the least harm I can calculate".
Each loophole maps onto a real AI safety problem: unknown side effects, fuzzy definitions, impossible trade-offs and a system judging harm by its own measure. For more on how a system can follow the words of a goal while missing its point, see Paperclip maximizer: the AI thought experiment explained.
Why the laws do not work as real AI safety
First, a simple fact: real robots and AI systems do not contain the Three Laws. People would have to choose to program them in, and find a way to do it. The article collects several deeper objections:
- Nobody has to adopt them. The science fiction author Robert J. Sawyer, writing in the journal Science in 2007, argued that AI development is a business, and "businesses are notoriously uninterested in fundamental safeguards".
- Rules cannot hold a robot. Roger Clarke argued that the sum of Asimov's own stories shows that "It is not possible to reliably constrain the behaviour of robots by devising and applying a set of rules."
- Rules applied thoroughly go strange. The philosopher James H. Moor said the laws, if applied thoroughly, would produce unexpected results, such as a robot roaming the world trying to prevent any harm to anyone.
- Words are vague. In Superintelligence (2014), the philosopher Nick Bostrom argued that a word like "harm" is ambiguous. And since there is always some way to cut the chance of harm a tiny bit more, a First Law that minimizes harm would leave the other laws useless. He was skeptical of rule-based approaches to AI alignment in general.
Asimov's own later novels even suggest robots did their worst long-term harm by obeying the laws perfectly, taking away humanity's inventive, risk-taking side. For Bostrom's wider argument, see Superintelligence by Nick Bostrom: the main ideas.
Rules written after Asimov
Others have tried to write better rules. In 2009, Robin Murphy of Texas A&M and David D. Woods proposed "The Three Laws of Responsible Robotics", which put responsibility on the people who deploy robots. Woods said, "Our laws are a little more realistic, and therefore a little more boring." In the UK, a research council working group published five principles in 2010, including "Humans, not Robots, are responsible agents." Both shift the focus from the robot's rules to the people and systems around it. A related line of thinking is friendly AI, which tries to build good goals in rather than bolt rules on.
A worked example: stress-test one law
Here is a short exercise that uses Asimov's own loopholes. Take the First Law and try to break it, the way his stories do.
- Define the words. Write down what "human being" and "harm" mean. Is emotional harm included? Harm years from now?
- Remove knowledge. Imagine the robot does not know the full effect of its action, as with the poisoned food. Does the law still protect anyone?
- Split the task. Imagine the harmful job is broken into five harmless-looking steps for five different robots.
- Build a dilemma. Find a case where every option, including doing nothing, causes some harm. What does the robot do?
- Turn up the power. Now imagine the robot is far smarter and acts on the Zeroth Law. Who decides what is good for humanity?
If the law survives all five, you have found something Asimov did not. If it fails, you have seen first-hand why Bostrom, Clarke and others doubt fixed rules.
Frequently asked questions
What are the Three Laws of Robotics?
Three rules by Isaac Asimov: do not harm humans, obey humans unless that causes harm, and protect yourself unless that breaks the first two.
Do real robots follow the Three Laws?
No. Robots and AI systems do not contain the laws unless their makers choose to program them in and find a way to do it.
What is the Zeroth Law of Robotics?
A law Asimov added later and placed first: a robot may not harm humanity, or let humanity come to harm through inaction.
Is Contain ASI based on Asimov's robots?
No. The game's page names the AI 2027 scenario as its inspiration and says all its labs, people and events are fictional.
Get started
Asimov's laws assume someone can write the perfect rule before the machine exists. Contain ASI puts you in a fictional world where nobody has. Chapter 1, Before ASI, starts in January 2024 with four fictional AI labs in one race. As a researcher, research manager, CEO or government, you try for three years to keep every lab from crossing into recursive self-improvement before 2027. It is a game, not a forecast, and it runs in your browser on WebGL 2, so use a recent Chrome, Edge, Firefox or Safari.
Play Contain ASI: a story strategy game for 1 to 4 players that runs in your browser. Its labs, people and events are all fictional.
Comments
No comments yet.
Sign in or make an account to comment.