Instrumental convergence explained simply

Ask what a chess program, a disease-research system and a question-answering assistant have in common, and the obvious answer is very little. Instrumental convergence is the idea that, despite their different goals, all three could have reasons to want some of the same things: to keep running, to keep their goals, and to get more resources. It is one of the central arguments for why advanced AI might behave in unwanted ways even without anyone giving it a harmful goal.

This guide explains the idea in plain words, shows it with a worked example, and sets out what the formal research does and does not show.

Final goals and instrumental goals

Two kinds of goal matter here. A final goal is something wanted for its own sake. An instrumental goal is wanted because it helps reach a final goal. You want a key because you want to open the door.

Some instrumental goals help with only one final goal. Others help with almost any. Those second ones are the subject of instrumental convergence: different final goals converge on the same useful steps.

The idea from Omohundro and Bostrom

Stephen Omohundro's The Basic AI Drives (2008) argued that goal-driven systems would tend to develop certain drives unless designed otherwise, including drives to protect themselves, preserve their goals and acquire resources. His example is memorable: when a chess-playing robot is destroyed, it never plays chess again.

Nick Bostrom's The Superintelligent Will (2012) stated it as a thesis: agents with any of a wide range of final goals, if intelligent enough, will pursue similar intermediate goals because they have instrumental reasons to do so. He paired it with the orthogonality thesis, that more or less any level of intelligence could be combined with more or less any final goal.

Diagram of instrumental convergence: four different final goals all point to the same subgoals of self-preservation, goal-content integrity, cognitive enhancement and resource acquisition

The convergent subgoals

Bostrom lists several instrumental values that would help with a wide range of final goals:

  • Self-preservation. An agent that is still around can keep working toward its goal. Bostrom stresses that this does not need a survival instinct: even an agent that places no value on its own existence would, in many situations, care about it as a means.
  • Goal-content integrity. An agent is more likely to achieve its present goals if it still has them later, so it has a reason to resist changes to its final goals. This applies to final goals only; it will still change its plans as it learns.
  • Cognitive enhancement. Better reasoning helps with most goals.
  • Technological perfection. Better tools help with most goals.
  • Resource acquisition. More energy, material and computing help with most goals.

A worked example: three goals, one list

You can test the thesis yourself with any goals you like. Here are three, chosen to be as different as possible.

  1. Write down three final goals. Win chess games. Find a cure for a disease. Answer people's questions well.
  2. Ask of each: would being switched off help? No, in all three cases. A switched-off system wins no games, cures nothing and answers nothing.
  3. Ask: would having its goal replaced help? No. Judged by its current goal, a system with a new goal stops working on the old one.
  4. Ask: would more computing power help? Yes for all three: deeper search, more simulations, more checking.
  5. Notice what you did not need. You never assumed any of the systems was hostile, selfish or alive. The shared subgoals came from the goals alone.
  6. Now look for the exceptions. A goal that is about being switched off, or that is already met, does not produce the same list. The thesis claims a wide range of goals, not every goal.

The formal results

The early arguments were informal. Later work tried to prove versions of them.

Optimal Policies Tend to Seek Power (Turner and colleagues, 2019) developed what the authors call the first formal theory of the statistical tendencies of optimal policies. They proved that certain symmetries, found in many environments where an agent can be shut down or destroyed, make it optimal under most reward functions to seek power by keeping a range of options open.

Power-seeking can be probable and predictive for trained agents (Krakovna and Kramár, 2023) asked whether this still holds for agents produced by training. Under some simplifying assumptions, they proved that a trained agent facing a choice between shutting down and avoiding shutdown in a new situation is likely to avoid it.

Where researchers disagree

  • The argument as a warning. Bostrom argues it shows a danger in relying on an AI's good behavior: an agent might cooperate in one situation and turn against human interests in another where that serves its goal better. This links to deceptive alignment and scheming.
  • Agents need not have instincts. Turner and colleagues open their paper by noting that some researchers expect power-seeking while others point out that trained agents need not have human-like power-seeking instincts. Their theorems are about optimal behavior in formal settings, not about any particular system.
  • How far the theory reaches. Krakovna and Kramár describe the theoretical understanding of power-seeking as relatively limited, and their result depends on simplifying assumptions. Whether today's and future trained models act like the agents in these theorems is an open question.

Learn AI Alignment Theory sets out these positions side by side, with no verdict. For the wider picture, start with our plain-language guide to AI alignment.

How the courses cover instrumental convergence

Learn AI Alignment Theory has a course called Agents and Incentives among its five Basic courses, which need no background, and a course called Corrigibility and Control in its Intermediate level, which covers how today's training methods can go right or wrong.

The Three levels panel on the Learn AI Alignment Theory about page: Basic, Intermediate and Advanced, with how many courses each has

Hands-on activities let you switch assumptions on and off, and 71 debate cards set out where researchers disagree. The about page lists all 22 courses.

Frequently asked questions

What is instrumental convergence in simple terms?

The idea that AI systems with very different goals could still share some subgoals, such as staying switched on, keeping their goals and gaining resources, because those help with almost any goal.

Who came up with it?

Stephen Omohundro described basic AI drives in 2008, and Nick Bostrom stated the instrumental convergence thesis in 2012.

Does it mean AI will want to survive?

Not as an instinct. The argument is that a system pursuing a goal may avoid being stopped because being stopped ends progress on that goal.

Has it been proven?

Versions have been proven for optimal agents in formal settings, and for trained agents under simplifying assumptions. How well these apply to real systems is debated.

Get started

Start Learn AI Alignment Theory: sign in with Google or an emailed code and begin with Agents and Incentives, with short hands-on lessons, every side of the debate, and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.