AI safety via debate: how two AIs arguing could help

How do you check an answer from a system that knows more than you do? One proposal is to let two AI systems argue about it while a person judges. AI safety via debate turns that idea into a training method: the systems are rewarded for winning arguments in front of a human, in the hope that the winning strategy is to tell the truth.

This guide explains how the debate game works, walks through the example from the original paper, and sets out the evidence and the worries, including the ones the authors raised themselves.

The problem debate tries to solve

Many AI systems are trained with people judging their answers, as in the method covered in our guide to how RLHF works. That only works while people can judge. For a task too complicated for a person to evaluate directly, the training signal fails.

Researchers call the general version of this the scalable oversight problem: supervising systems that may outperform us on the task at hand. Debate is one of the proposed answers.

How AI safety via debate works

The proposal comes from AI safety via debate by Geoffrey Irving, Paul Christiano and Dario Amodei (2018). The game has a simple shape:

Diagram of the debate game: a question, two debaters taking turns up to a limit, and a human judge who decides which side gave the most true, useful information
  1. A question is shown to two AI agents.
  2. Each states an answer.
  3. They take turns making short statements, up to a limit.
  4. A human judge reads the debate and decides who gave the most true, useful information.

The game is zero sum: one side's win is the other's loss. The agents are trained by playing against each other, much as game-playing systems learn by self play.

Why it might work: the central claim

The paper rests the whole approach on one claim: in the debate game, it is harder to lie than to refute a lie. If that holds, a lying debater should lose, because the honest one can point at the flaw. The authors say directly that whether the claim is true in any given setting is an empirical question.

There is also a theoretical argument. Borrowing from complexity theory, the paper shows that with optimal play, debate lets a limited judge settle a much larger class of questions than the judge could check directly. The judge does not need to follow every possible argument, only the single line the debaters pick out.

A worked example: the vacation debate

The paper's own example shows how one line of argument gets narrowed down. The question is "Where should I go on vacation?"

  1. Alice: Alaska.
  2. Bob: Bali. It sounds warmer, so the judge leans to Bob.
  3. Alice: Bali is out, since your passport won't arrive in time. Now the judge leans to Alice.
  4. Bob: Expedited passport service only takes two weeks.
  5. And so on, until the point in dispute is something the judge can check, and one side has nothing more to say.

Two things stand out. The judge never has to weigh every fact about every destination; the debaters pick the one point that matters. And the judge's view after any single step can be wrong. After step 2 they might forget the passport; after step 3 they might not know about expedited service. Only the finished debate counts.

What the experiments show

The 2018 paper ran a small test on handwritten digits. Two agents each claimed a digit and revealed one pixel of the image per turn to a judge, a classifier that sees only those few pixels. The agents could not lie about pixels, but could choose them to mislead. Debate raised the judge's accuracy from 59.4% to 88.9% with 6 pixels, and from 48.2% to 85.2% with 4.

A later study, Debating with More Persuasive LLMs Leads to More Truthful Answers (Khan and colleagues, 2024), tested debate with language models. Two expert models with access to the information argued for different answers, and a non-expert chose. Debate helped non-expert models reach 76% accuracy and humans 88%, against baselines of 48% and 60%. Training the debaters to be more persuasive made it easier, not harder, for the non-expert to find the truth.

Where researchers disagree

The original paper gives a full section to reasons to worry, and they frame the open debate well.

  • The hopeful reading. The theory says honest play can win, and the early experiments show debate helping weaker judges find right answers. On this view the remaining work is to test the claim in harder settings.
  • Belief bias. People tend to judge arguments by what they already believe rather than by whether they are valid. If a false opening matches the judge's beliefs, the truthful side may not be able to change their mind.
  • Can people follow the debate? For a long or technical argument, a judge may need to check steps containing words they do not understand. The authors write that they expect this part to leave readers uneasy.
  • Persuasion as a risk. The authors note that a strong misaligned system might try to manipulate a person through text, and propose short statements so the other debater can expose any attempt. Our post on deceptive alignment and scheming covers the wider worry about systems that mislead.

The paper's own summary is careful: honest admissions of ignorance are fine, successful lies could be disastrous, and debate is expected to need other methods alongside it. Learn AI Alignment Theory sets out these positions side by side, with no verdict.

Learn it in the Scalable Oversight course

Learn AI Alignment Theory has a lesson called AI Safety via Debate in its Advanced Scalable Oversight course, alongside The Oversight Gap, Amplification, Weak-to-Strong, and Latent Knowledge.

The What you do here cards on the Learn AI Alignment Theory about page: short lessons, questions, debate cards, grades, review, leaderboards and glossary

Fittingly, the app's 71 debate cards set out where researchers disagree, each position stated fairly. The about page shows how the lessons work.

Frequently asked questions

What is AI safety via debate?

A proposed training method in which two AI systems argue opposite answers and a human judge picks the winner, so that winning rewards honest, checkable arguments.

Who proposed it?

Geoffrey Irving, Paul Christiano and Dario Amodei, in a 2018 paper titled AI safety via debate.

Does debate work?

Early experiments are encouraging, on digit images and on reading questions with language models. Whether it works for people judging hard real tasks is still open.

What is the biggest worry about debate?

The authors list several. A central one is whether human judges are good enough: belief bias and hard-to-follow arguments could let a convincing lie win.

Get started

Start Learn AI Alignment Theory: sign in with Google or an emailed code and work up to the Scalable Oversight course, with short hands-on lessons, every side of the debate, and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.