Corrigibility: why an AI that accepts correction is hard to build

Every machine should have an off switch. For a simple machine that is easy: it has no opinion about being switched off. Corrigibility asks how to keep that property in an AI system that pursues goals, because a system that is trying to achieve something has a reason to avoid being stopped or changed before it gets there.

This guide explains what corrigibility means, why the off switch turned out to be a real research problem, and the main ideas researchers have tried.

What corrigibility means

The term was set out in Corrigibility by Nate Soares, Benja Fallenstein, Eliezer Yudkowsky and Stuart Armstrong (2015). They call an AI system corrigible if it cooperates with what its creators regard as a corrective intervention, even though rational agents by default have incentives to resist being shut down or having their preferences changed.

The paper lists what that should include. A corrigible system:

  • tolerates, and preferably helps with, attempts to change it or turn it off;
  • does not try to manipulate or deceive the people overseeing it;
  • tends to repair broken safety measures, such as a broken shutdown button, or at least reports them;
  • keeps people able to correct or shut it down, even when it builds new subsystems or changes itself.

Why a goal-driven AI might resist

The worry is not that an AI would want to survive the way an animal does. Stephen Omohundro's The Basic AI Drives (2008) put it in one line: when a chess-playing robot is destroyed, it never plays chess again. Whatever the goal, being switched off stops progress on it.

The same logic applies to having its goal changed. Soares and colleagues note that an agent's current goal is usually better served if the agent keeps pursuing it, so by default it has a reason to stop its programmers from changing it, even if it knows they meant to give it a different goal. Our plain-language guide to AI alignment shows where this fits in the wider problem.

The shutdown button problem

The 2015 paper focuses on the simplest case: a shutdown button. It asks for a goal that makes an agent shut down when the button is pressed, while giving it no reason to stop the press, no reason to cause the press, and no reason to build helpers that ignore the button. The authors write that they know of no solution meeting all of these.

Diagram of the off-switch game: a certain robot has an incentive to disable its switch, an uncertain robot wants to keep it, and the general shutdown problem is still open

A worked example: pricing the off switch

The paper's central difficulty can be shown with simple numbers. The values below are made up for illustration.

  1. The normal goal. A robot expects 10 points of value from finishing its task.
  2. Add a shutdown reward. If the button is pressed, the robot gets a fixed reward for shutting down. The question is how big to make it.
  3. Too low. Set it at 2. Now a press costs the robot 8 points, so it has a reason to stop people from pressing: block the button, or persuade them not to.
  4. Too high. Set it at 20. Now a press gains it 10 points, so it has a reason to get the button pressed: frighten people, or press it itself.
  5. Exactly equal. Set it at 10, so the robot does not care either way. The paper's name for this idea is utility indifference. But a robot that does not care whether the button works also will not spend any effort keeping it working. The paper shows such an agent would pay no cost at all to make new machinery it builds respond to shutdown.

The authors add that this balance is hard to keep, because the robot's expected value from its task keeps changing as it learns.

The off-switch game: uncertainty as a way out

The Off-Switch Game by Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell (2016) takes a different angle. A robot can act, switch itself off, or disable its off switch; a human can press the switch if it still works.

A robot that treats its goal as certain has a reason to disable the switch, unless the human is perfectly rational. But a robot that is uncertain about what outcome is actually good, and treats the human's choice to press as evidence about it, wants to keep the switch. The authors conclude that giving machines an appropriate level of uncertainty about their objectives leads to safer designs.

Where researchers disagree

  • Corrigibility as an open formal problem. Soares and colleagues found that every proposal they examined, including utility indifference, failed at least one of their requirements, and they describe even the simple shutdown problem as wide open.
  • Uncertainty as the route. Hadfield-Menell and colleagues argue that the incentive to disable a switch comes from an agent being sure of its goal, and that building in the right uncertainty removes it in their model. They present this as a useful generalization of the classical idea of a rational agent.
  • What remains open. The off-switch result depends on the robot treating the human's choices as informative. How these models carry over to large trained systems, which are not built as explicit utility maximizers, is a live question.

Learn AI Alignment Theory sets out these positions side by side, with no verdict. For the related worry that a system might only appear to accept correction, see our post on deceptive alignment and scheming.

How the courses cover corrigibility

Learn AI Alignment Theory has a Corrigibility and Control course in its Intermediate level, which covers how today's training methods can go right or wrong. Agents and Incentives, one of the five Basic courses that need no background, is a gentler place to begin.

The Learn AI Alignment Theory home page with its title, a short description and the buttons to sign in with Google or an emailed code

Lessons include hands-on activities where you switch assumptions on and off, and 71 debate cards set out where researchers disagree. The about page lists all 22 courses.

Frequently asked questions

What does corrigibility mean in AI?

An AI system is corrigible if it cooperates with being corrected, changed or shut down by the people overseeing it, even when its own goals would give it a reason to resist.

Why not just add an off switch?

A system pursuing a goal may have a reason to stop the switch being used, or to get it pressed, depending on how it is rewarded. Making it neutral without making it careless is the hard part.

What is utility indifference?

A proposal to make an agent value being shut down exactly as much as carrying on. The 2015 paper shows it leaves the agent with no reason to keep shutdown working.

Is corrigibility solved?

No. The paper that named it calls even the simple shutdown problem wide open, and later work offers partial answers in simplified models.

Get started

Start Learn AI Alignment Theory: sign in with Google or an emailed code, then take Agents and Incentives and Corrigibility and Control, with short hands-on lessons, every side of the debate, and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.