Coherence Theorems and AI Goals Explained: The Case and the Critique

If you have heard someone say that a powerful AI will "act like a maximizer", they were probably leaning on coherence arguments. This post gives you coherence theorems and AI goals explained in plain terms: what the argument claims, a money pump worked through step by step, why it matters for alignment, and the 2023 critique that says the theorems people cite do not say what they are taken to say.

Coherence theorems and AI goals: the claim in plain words

The claim, as EJT summarizes it in "There are no coherence theorems" (2023), is that there are theorems showing that unless an agent can be represented as maximizing expected utility, it is liable to pursue strategies that are dominated by some other available strategy.

A strategy is dominated, in EJT's definition, when another strategy is worse in no respect the agent cares about and better in at least one. Taking a dominated strategy means losing something you value for nothing. Authors describe the same idea as "money-pumping", "Dutch books" or "shooting oneself in the foot".

Eliezer Yudkowsky's essay "Coherent decisions imply consistent utilities", written for Arbital in 2017, makes the positive case: if you avoid dominated strategies, you must behave as if you weigh outcomes by consistent numbers, which we then call utilities.

Diagram of the coherence argument: circular preferences lead to wasted taxi rides, the claim that unexploitable agents act as expected utility maximizers, and EJT's 2023 challenge about completeness

The coherence argument, and the 2023 challenge to it.

A money pump, step by step

Yudkowsky's own example is about places: if you prefer Berkeley to San Francisco, San Jose to Berkeley, and San Francisco to San Jose, "you're going to waste a lot of time on taxi rides". Here is the same idea with prices, as an illustration you can follow.

An agent has circular preferences over three fruits. It prefers a pear to a banana, a banana to a cherry, and a cherry to a pear. It starts holding a cherry.

  1. A trader offers a banana for the cherry plus 1 cent. The agent prefers the banana, so it pays.
  2. The trader offers a pear for the banana plus 1 cent. The agent prefers the pear, so it pays.
  3. The trader offers a cherry for the pear plus 1 cent. The agent prefers the cherry, so it pays.

The agent is back where it started, holding a cherry, 3 cents poorer, and the trader can repeat the loop. Each step looked like an improvement; the whole cycle was a loss. That is the intuition coherence arguments build on: inconsistency is a leak that someone can exploit.

From preferences to expected utility

The best-known result in this area is the von Neumann and Morgenstern (VNM) theorem. EJT lists its four axioms, here in plain words:

  • Completeness: for any two options, you prefer one or rate them equally.
  • Transitivity: if you rate A at least as high as B, and B at least as high as C, you rate A at least as high as C.
  • Independence: mixing the same third option into two gambles does not flip which one you prefer.
  • Continuity: if you rank A above B and B above C, some gamble between A and C is worth exactly as much as B.

If your preferences over gambles meet these, your choices can be described as maximizing expected utility. A worked example: say an agent's utilities are 0 for nothing, 60 for a sandwich and 100 for a full dinner. A sure sandwich is worth 60. A 50 percent chance of dinner is worth 0.5 times 100, which is 50, so the agent takes the sandwich. Raise the chance to 70 percent and the gamble is worth 70, so it switches. One consistent scale ranks every gamble.

Note the words "described as". The theorem is about how choices can be summarized, not about a number stored inside the agent.

Why coherence arguments matter for AI goals

EJT writes that coherence arguments "seem to be a moderately important part of the basic case for existential risk from AI". The argument runs roughly:

  1. Capable AI systems will avoid being exploited, because exploitation wastes what they value.
  2. Agents that avoid dominated strategies act like expected utility maximizers.
  3. A maximizer of almost any goal benefits from more resources and from not being switched off, the idea known as instrumental convergence.
  4. So capable AI will tend to pursue its goals hard, and slightly wrong goals become dangerous.

Its appeal is that it does not depend on any one design: pressure toward competence alone is supposed to push systems toward goal-directed behavior.

The critique: "There are no coherence theorems"

EJT, writing as part of the CAIS Philosophy Fellowship, argues that step 2 is not a theorem. In his words: "There are no coherence theorems. Authors in the AI safety community should stop suggesting that there are. There are money-pump arguments, but the conclusions of these arguments are not theorems."

His main target is Completeness. The VNM theorem assumes it; it does not prove that an agent who avoids money pumps must have it. EJT describes a policy that an agent with incomplete preferences can follow, and argues that such an agent is immune to all possible money pumps for Completeness. If so, an agent can avoid exploitation without being representable as an expected utility maximizer.

Where the debate stands

Both posts are on LessWrong and both sides are argued seriously.

  • The coherence side: avoiding dominated strategies pushes an agent toward behaving as if it has consistent utilities, and that is the behavior safety work worries about.
  • The critique: the cited theorems take Completeness as an assumption, money-pump arguments for it rest on doubtful assumptions, and an agent with incomplete preferences can avoid being pumped.

A useful habit either way: when someone claims what capable AI "will" do, list the assumptions and switch each off, as in the money pump above. If the agent never has to rank two options, does the conclusion survive?

If coherence does push toward maximizing, one response is to design agents that do not maximize: see quantilizers and mild optimization. For the wider field that studies these questions, see agent foundations.

Frequently asked questions

Do coherence theorems prove AI will have goals?

No. The VNM theorem shows that preferences meeting its four axioms can be described as maximizing expected utility. Whether capable AI systems will meet those axioms is the debated part.

What is a money pump?

A sequence of trades, each of which an agent with inconsistent preferences accepts, that leaves it worse off than where it started, as in the fruit example above.

Why does completeness matter so much?

Because the VNM theorem assumes it. EJT argues that an agent with incomplete preferences can avoid money pumps, so avoiding exploitation does not force an agent to be an expected utility maximizer.

How do coherence arguments relate to instrumental convergence?

Coherence arguments are meant to show capable agents act like maximizers. Instrumental convergence then says maximizers of many goals share sub-goals like gaining resources and avoiding shutdown.

Get started

On Learn AI Alignment Theory, the Basic course Agents and Incentives and the Advanced courses Agent Foundations and The Big Debates are the ones whose titles match this topic. Lessons take about 8 minutes, hands-on activities include sliders and scenarios where you switch assumptions on and off, and debate cards set out each serious position with no verdict. Every lesson lists its sources and separates what is known from what is still open. You sign in with Google or an emailed code. Read more on the about page.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.