The Off-Switch Game Explained Simply, With Numbers
Can you build an AI that lets you switch it off? The off-switch game is a small model that asks exactly that, and its answer is surprising: a robot that is sure of its goal has a reason to block the switch, and a robot that is unsure has a reason to leave it alone. This post gives you the off-switch game explained simply, with a worked example you can check by hand, and the later papers that test where it breaks.
The off-switch game explained in one paragraph
In 2016, Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell published The Off-Switch Game. It is a game between a human, H, and a robot, R. H can press R's off switch. R can disable it. The paper asks when R has a reason to let H keep that power.
The starting worry is in the paper's opening: many rational agents have "strong incentives for self-preservation." Not from an instinct, but because an agent that maximizes expected utility "cannot achieve whatever objective it has been given if it is dead." This links to instrumental convergence.
The setup: a robot, a human and a button
R moves first and has three choices:
- a: act now, bypassing the human. Disabling the off switch is one way. The outcome is worth Ua to the human.
- w(a): tell H it wants to do a, and wait. H can then switch R off, or let it go ahead.
- s: switch itself off. This outcome is worth 0.
The paper says the model is meant to be "as simple as possible." It leaves out a cost for disabling the switch, extra options for R, a second human, and other complications, and lists them as future work.
A worked example: why uncertainty makes the robot wait
The numbers here are made up; the logic is the paper's. Say R wants to do something, and it is not sure whether the human will like it:
- 60% chance the outcome is worth
+10to the human. - 40% chance it is worth
-10.
Now score each choice from R's point of view, where R's goal is whatever is good for H:
- a (act now):
0.6 × 10 + 0.4 × (-10) = 2. - s (switch off):
0. - w(a) (ask and wait): suppose H knows which case it is and acts rationally. H lets R go ahead when the outcome is +10 and switches it off when it is -10. So the value is
0.6 × 10 + 0.4 × 0 = 6.
Waiting wins, 6 against 2. The human's choice carries information R lacks, so leaving the switch alone is worth something to R. That is the paper's key insight: for R to want to keep its off switch, "it needs to be uncertain about the utility" and treat H's actions as evidence about it.

Why a certain robot has a reason to block the switch
Now make R sure. It believes the outcome is worth +10, full stop.
- a: 10.
- w(a): if H is perfectly rational and also sees +10, H lets R go ahead: 10. A tie.
- w(a) with a human who sometimes errs: if H switches R off by mistake some of the time, waiting is worth less than 10, and acting now wins.
This matches the paper's result: an agent that takes its reward function for granted has an incentive to disable the off switch, "except in the special case where H is perfectly rational." Certainty plus an imperfect human is the dangerous mix.
The paper also shows why a human treated as pure noise does not help. If H's choice is random, unrelated to Ua, waiting is just a blend of Ua and 0, which is never better than the best of the two. R gains nothing by asking someone whose answer carries no information.
When the argument breaks: the follow-up papers
How rational is the human?
Wängberg and colleagues (2017) point out that the original analysis "is not fully game-theoretic," because it models the human as irrational and computes the robot's best action under "unrealistic normality and soft-max assumptions." They redo it with the human as a rational player with a random utility function, which lets them compute the robot's best action under any belief and any level of human error.
What if the robot's model is wrong?
Ryan Carey's "Incorrigibility in the CIRL Framework" (2017) presses on the other assumption. The incentive to obey shutdown holds if the shutdown command carries information about what is valuable. But that "is not robust to model mis-specification," such as programmer errors. Carey shows cases where errors in the reward model remove the incentive to follow shutdown commands. He argues for systems that obey shutdown under weaker assumptions, for example that "one small verified module is correctly implemented."
What the off-switch game means for real oversight
The game turns a vague hope, "we can always switch it off," into a precise question about incentives. It connects to corrigibility, the property of a system that accepts correction.
Researchers draw different lessons. The original authors conclude that giving machines "an appropriate level of uncertainty about their objectives leads to safer designs." Carey's results suggest that uncertainty is only as good as the model it lives in, and that shutdown should not depend on getting the whole model right. Both lines of work use the same simple game, which is why it is a good place to start.
Frequently asked questions
Who came up with the off-switch game?
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell, in a 2016 paper titled The Off-Switch Game.
Is the off-switch game the same as corrigibility?
No. Corrigibility is the broader property of accepting correction. The off-switch game is one formal model of one part of it: whether a robot lets a human switch it off.
Does uncertainty solve the shutdown problem?
In the original model it gives the robot a reason to defer, if the human is rational enough. Carey shows that errors in the robot's model can remove that reason, so the question stays open.
Do I need advanced maths to follow it?
No. If you can follow the worked example above, multiplying chances by values, you have the core of the argument.
Get started: learn AI alignment theory step by step
On Learn AI Alignment Theory, Agents and Incentives (Basic) covers why goals create incentives, Corrigibility and Control (Intermediate) takes on shutdown and correction, and Agent Foundations and The Big Debates (Advanced) go deeper. Lessons run about 8 minutes, and hands-on activities let you switch assumptions on and off, much like changing the human's rationality in this game. Every lesson lists its sources, and debate cards set out each serious position with no verdict. You sign in with Google or an emailed code.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.