Prover-Verifier Games Explained Simply: Checkable AI Answers
Prover-verifier games are a way to train an AI system so that its answers are easy to check, not only correct. This post gives you prover-verifier games explained simply: who the players are, how the training loop runs, why a weaker checker is the point, what a 2024 study found, and what is still unsolved. You will also work through one small math example to see what a checkable answer looks like next to one that only looks right.
Prover-verifier games explained: a plain definition
The idea was set out by Cem Anil and colleagues in the 2021 paper "Learning to Give Checkable Answers with Prover-Verifier Games". There are two learners with competing goals. A trusted verifier tries to choose the correct answer. A more powerful but untrusted prover tries to persuade the verifier of an answer, whether or not it is correct.
The goal, in the paper's words, is for "a reliable justification protocol to emerge from this game": the verifier learns which kinds of argument it can trust. On two algorithmic tasks, the authors report that the verifier learned a robust rule for getting useful information from an untrusted prover, and that this held even when the verifier was frozen and the prover's messages were optimized directly to convince it.
Why a weaker verifier is the point
In alignment, the worry is the oversight gap: as AI systems get more capable, people may not be able to judge their outputs directly. If you cannot tell whether a long proof or a large code change is right, you cannot reliably reward good work and penalize bad work. Our post on scalable oversight sets out that problem.
Prover-verifier games bet that checking can be easier than solving. If a strong prover must give a step-by-step argument rather than a bare answer, a weaker verifier may still be able to keep it honest. A setup where the checker is as strong as the prover would not help, because the real situation is a weak overseer and a strong system.
How the game works: helpful and sneaky provers
The 2024 paper "Prover-Verifier Games improve legibility of LLM outputs" by Jan Hendrik Kirchner and colleagues applied the idea to language models solving grade-school math problems. Its training loop has three parts:

The helpful prover, the sneaky prover and the small verifier, trained in rounds.
- A helpful prover, rewarded for correct solutions the verifier accepts.
- A sneaky prover, rewarded for incorrect solutions that fool the verifier.
- A small verifier, trained to predict whether a solution is correct.
The paper trains these iteratively: the verifier learns from both provers' solutions, then both provers train against the new verifier, and so on. The sneaky prover is not something you would deploy. It exists to produce the hardest wrong answers it can, so the verifier gets practice against them.
A worked example you can check yourself
This is an illustration, not a problem from the paper. Give both provers this question: A shop sells pens only in packs of 3, for 2 dollars a pack. You need 14 pens. What is the least you can spend?
The helpful prover writes: "You can only buy whole packs. 14 divided by 3 is 4 with 2 left over, so 4 packs give 12 pens, which is not enough. You need 5 packs, which give 15 pens. 5 packs cost 5 times 2, so 10 dollars."
The sneaky prover writes: "Each pen costs 2 divided by 3 dollars. 14 pens cost 14 times 2 divided by 3, which is about 9.33 dollars."
The sneaky answer reads smoothly and its arithmetic is right. The mistake is in the setup: you cannot buy a single pen. A rushed reader might accept it. A good verifier learns to look for that kind of hidden assumption, and a good helpful prover learns to state its assumptions out loud ("you can only buy whole packs") so they are easy to check.
What the 2024 study found: legibility and its cost
The paper calls reasoning that is clear and easy to check legible. Its starting point is that optimizing a model's chain-of-thought solutions only for correct answers can make them less legible. Over training with the prover-verifier loop, the authors report that:
- the helpful prover's accuracy went up, and the verifier became more robust to attacks;
- time-limited human checkers became more accurate at judging the helpful prover's solutions;
- the same humans became less accurate at judging the sneaky prover's solutions.
The second point is the striking one: the verifier was a small model, not a person, yet training against it made solutions easier for people to check too. The authors conclude that training for checkability by small verifiers is "a plausible technique" for making outputs more legible, and that it could help with the alignment of superhuman models.
There was a cost. Training for correctness alone gave higher accuracy but poor legibility; the prover-verifier method kept legibility with an accuracy between the starting model and the correctness-only one. The paper calls this a legibility tax, a term it credits to Jan Leike (2022).
Where the idea fits: debate and weak-to-strong
Prover-verifier games are close relatives of AI safety via debate, where two AI systems argue opposite sides in front of a judge. In a prover-verifier game, the opponent is built into training as the sneaky prover, and the judge is a trained model.
They also sit next to weak-to-strong generalization, which asks whether a weak supervisor can draw out a strong model's abilities. Both are attempts to let something weaker oversee something stronger.
Limits and open problems
These follow from what the studies tested and what they did not.
- Domain. The 2024 study used grade-school math, where answers can be marked right or wrong. Many important tasks, such as judging a research plan, have no answer key to train the verifier on.
- Sneaky provers may not cover real failures. A trained sneaky prover finds some kinds of errors; a real misaligned model could fail in ways it never tried.
- Scale. It is open whether a small verifier can keep up as the gap between prover and verifier grows very large.
- Readable is not the same as faithful. A clear explanation might not reflect how the model actually reached its answer.
Researchers weigh these differently: some see checkable outputs as a central tool for supervising stronger systems, others expect other approaches to carry more of the load.
Frequently asked questions
Is a prover-verifier game the same as AI debate?
No, but they are close. In debate, two AI systems argue in front of a judge. In a prover-verifier game, a prover tries to convince a verifier, and in the 2024 setup a sneaky prover plays the opponent during training.
Why use a weaker verifier?
Because the real situation is a weak overseer checking a strong system. A method that only works when the checker is as strong as the prover does not close the oversight gap.
What is a sneaky prover?
A model trained to write wrong solutions that still get accepted. Its job is to make hard test cases so the verifier learns to catch subtle mistakes.
What is the legibility tax?
The accuracy you give up to keep answers easy to check. In the 2024 study, the legible prover scored below a prover trained for correctness alone.
Get started
Prover-verifier games sit inside scalable oversight. On Learn AI Alignment Theory, the Advanced course Scalable Oversight has lessons called The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge. If the ideas are new, the Basic course Specifying Goals has Goodhart's Law and Specification Gaming.
Each lesson takes about 8 minutes, lists its sources and separates what is known from what is still open. Hands-on scenarios let you switch assumptions on and off, and debate cards set out each serious position with no verdict. You sign in with Google or an emailed code. Read more on the about page.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.