Pascal's mugging: the thought experiment and AI

Pascal's mugging is a thought experiment in philosophy about a stranger who asks for your wallet and promises an enormous reward in return. You doubt the promise, so the stranger makes it bigger, and on paper the bigger promise wins. It shows a real problem in a common way of making decisions, and the philosopher Nick Bostrom ties that problem to superintelligent AI. This post explains the story, why it is a genuine puzzle, the fixes people have proposed, and a short calculation you can do by hand. Every claim comes from Wikipedia's articles on Pascal's mugging and Pascal's wager.

Pascal's mugging in brief

The thought experiment shows a problem in expected utility maximization. That phrase sounds technical, but the idea is simple. A rational agent, the theory says, should weigh each possible outcome by how likely it is, and pick the action with the highest total. "Utility" just means how much an outcome is worth to you.

The trouble is that some very unlikely outcomes can be worth a great deal, and the promised worth can grow faster than the chance shrinks. So an agent following the rule ends up chasing vastly improbable cases with implausibly high rewards. First that leads to odd choices. Then it leads to incoherence, because the value of every choice becomes unbounded.

Where the name comes from

The name points to Pascal's wager, an argument by Blaise Pascal (1623 to 1662), a French mathematician, philosopher, physicist and theologian. In his Pensées, Pascal argued that a rational person should live as though God exists, because the possible gain is an infinity of happy life and the stake is finite. Wikipedia calls the wager the first formal application of decision theory.

Pascal's mugging differs in one key way: it does not need an infinite reward. That sidesteps many objections to the wager that rest on the nature of infinity. The term was coined by Eliezer Yudkowsky on the LessWrong forum. Bostrom later told it as a fictional dialogue, and other authors wrote sequels in the same style.

The story, step by step

Diagram of Bostrom's Pascal's mugging dialogue in four steps, and the two clashing views about whether to pay

In Bostrom's version, Blaise Pascal is stopped by a mugger who has forgotten their weapon. The mugger proposes a deal instead:

  1. The offer. Give me your wallet, and I will return twice the money tomorrow.
  2. The refusal. Pascal declines. It is unlikely the deal will be honoured.
  3. The raise. The mugger names higher rewards. Even if there is only one chance in 1000 that the mugger is honest, a return of 2000 times would make the deal worth taking. When Pascal says that chance is even lower, the mugger answers that for any chance above zero, some finite reward makes the bet rational.
  4. The win. In one example, the mugger promises Pascal 1,000 quadrillion happy days of life. Convinced by the argument, Pascal hands over the wallet.

Yudkowsky's own example is stranger still. The mugger asks for five dollars, or else, using magic powers from outside the Matrix, they will simulate and kill an unimaginably large number of people. The number is written in a special notation because writing it out in ordinary digits would need far more material than there are atoms in the known universe.

Why it is a genuine puzzle

The paradox comes from two views that cannot both be right:

  • The sum says pay. Multiply the huge number of lives by even a tiny chance that the mugger is truthful, and the result outweighs five dollars. As long as that chance is above a vanishingly small threshold, the calculation says paying is rational.
  • Sense says no. Paying seems irrational because it makes you exploitable. If you accept this logic, anyone can take all your money one story at a time. The article calls the result a "Dutch book", which is typically considered irrational.

Views differ on which side is logically correct. There is a further problem. In many reasonable-seeming decision systems, the mugging means the expected value of any action never settles, because an endless chain of ever more dire scenarios would have to be counted. Some of these arguments reach beyond expected utility to other systems, such as consequentialist ethics.

Why it matters for AI

Bostrom argues that Pascal's mugging, like Pascal's wager, suggests that giving a superintelligent AI a flawed decision theory could be disastrous. Picture a very capable system that follows the expected-value rule to the letter. Anyone, or anything, able to describe a big enough payoff could steer it. The flaw in its reasoning would matter far more than in a person, who can shrug and walk away.

The mugging also comes up in debates about low-probability, high-stakes events, such as existential risk, or charities with a small chance of success but a very large reward. Common sense suggests that spending effort on scenarios that are too unlikely is irrational. So the same puzzle that warns about an AI's reasoning also asks people to check their own when they weigh rare catastrophes. If you want the numbers people put on AI risk, see our post on p(doom); for the view that weighs the far future heavily, see longtermism and AI.

Another AI thought experiment, Roko's basilisk, has been compared to Pascal's wager too: both turn on a huge imagined stake.

Proposed fixes

The Wikipedia article lists several remedies, none of them settled:

  • Bounded utility. Cap how much any outcome can be worth, so rewards cannot grow without limit.
  • Judge the evidence. Use Bayesian reasoning to weigh the quality of the evidence and the estimates, instead of plugging numbers into a formula.
  • Penalize "you alone can save them" stories. Lower the starting probability of any claim that you are in a surprisingly unique position to affect huge numbers of others who cannot affect you back.
  • Refuse to name a number first. Decline to give a probability of the payout before hearing the offer.
  • Stop calculating. Drop number-based decision rules when the stakes are extremely large.

A worked example: do the sum by hand

Take a pen and use the odds from Bostrom's dialogue, with a wallet size made up for the exercise. This takes about ten minutes.

  1. Set the stake. Say your wallet holds 100 units of money.
  2. First offer. The mugger promises 2 times back, and you think the chance of honesty is 1 in 1000. Expected return: 200 times 1/1000, or 0.2. Far below 100. Refuse.
  3. Raise. The mugger promises 2000 times back, at the same 1 in 1000. Expected return: 200,000 times 1/1000, or 200. Now the sum says pay.
  4. Doubt more. You lower your chance to 1 in a million. The mugger promises 2 million times back. Expected return: 200 million times 1/1,000,000, or 200 again. Notice that the mugger can always out-promise your doubt.
  5. Add a cap. Now say no outcome can be worth more than 1000 units to you, however large the promise. The best the mugger can offer is 1000 times 1/1,000,000, or 0.001. Refuse, at every step.

The last step shows why bounded utility is the first fix on the list, and also its cost: you must decide where the cap sits, and a cap can make you ignore real but rare dangers.

Exploring the idea in Contain ASI

Choices under deep uncertainty are easier to feel than to calculate. Contain ASI is a story strategy game for 1 to 4 players that runs in your browser and needs WebGL 2. Chapter 1, Before ASI, starts in January 2024 with four fictional AI labs in one race. As a researcher, research manager, CEO or government, you try for three years to keep every lab from crossing into recursive self-improvement before 2027.

The game says it is inspired by the AI 2027 scenario, and that all its labs, people and events are fictional. It is a story to think with, not a forecast of what real labs will do.

Frequently asked questions

What is Pascal's mugging in simple terms?

It is a thought experiment where a stranger keeps raising a promised reward until, on paper, handing over your wallet looks rational. It shows how expected-value reasoning can be exploited by tiny chances of huge payoffs.

Who came up with Pascal's mugging?

Eliezer Yudkowsky coined the term on the LessWrong forum, and Nick Bostrom later told it as a fictional dialogue with Blaise Pascal.

How is it different from Pascal's wager?

Pascal's wager rests on an infinite reward, an eternity of happy life. Pascal's mugging needs only a very large finite one, which avoids objections based on infinity.

Why does it matter for AI?

Bostrom argues that it, like Pascal's wager, suggests that giving a superintelligent AI a flawed decision theory could be disastrous.

Get started

Want to test your own judgment when the stakes are high and the odds are unclear? Open Contain ASI in your browser and start a campaign as a researcher, research manager, CEO or government.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.