Weak-to-strong generalization: can weak AI supervise strong AI?
Weak-to-strong generalization is the question of whether a weak supervisor can train a model that ends up stronger than the supervisor itself. It matters for AI safety because people may one day need to supervise systems that are better than them at a task. Researchers study it today by letting a small AI model play the part of the human.
This guide explains the idea, the way it is measured, what the main experiments found, and why researchers read the results differently.
The problem it is meant to study
Much of today's AI training relies on people judging model behavior. Our guide to how RLHF works explains how those judgments become a training signal. The method assumes the judge can tell a good answer from a bad one.
Collin Burns and colleagues at OpenAI start their 2023 paper, Weak-to-Strong Generalization, from where that assumption may fail. Future superhuman models, they write, will behave in ways too complex for humans to reliably evaluate. Humans will only be able to supervise them weakly.
This is one form of the wider scalable oversight problem: keeping human checks useful when the system knows more than the checker.
What weak-to-strong generalization means
We cannot yet test supervision of a superhuman model, so Burns and colleagues study an analogy. A small model stands in for the human. A much larger model stands in for the future superhuman system. The question becomes: if the large model is trained only on labels from the small one, does it learn just the small model's answers, mistakes included, or does it do better?
If it does better, it has generalized beyond its supervisor. That is weak-to-strong generalization. The paper tested this with language models from the GPT-4 family, on language tasks, chess puzzles and reward modeling.
How it is measured
The paper compares three models on the same task:
- The weak supervisor: a small model trained on the true labels.
- The weak-to-strong student: a large model trained only on labels the weak supervisor produced.
- The strong ceiling: the same large model trained on the true labels, showing what it could do with perfect supervision.
From these, the authors define the performance gap recovered, or PGR: the share of the gap between the weak supervisor and the strong ceiling that the student recovers. A PGR of 1 means the student reached the ceiling. A PGR of 0 means it did no better than its supervisor.

A worked example: computing the gap yourself
You can follow the paper's arithmetic with simple numbers. The numbers below are made up for illustration.
- Train the weak supervisor. A small model learns to sort film reviews into positive and negative from true labels. On a test set it scores 60 percent.
- Let it label new data. It labels a fresh batch of reviews. Some of its labels are wrong, and nobody corrects them.
- Train the student on those labels. A much larger model learns only from the weak labels. On the same test set it scores 70 percent.
- Find the ceiling. Train the large model on true labels instead. It scores 80 percent.
- Compute the PGR. The student gained 10 points over its supervisor, out of a possible 20. So PGR = 10 ÷ 20 = 0.5. The student recovered half the gap.
- Check the mistakes. Look at the reviews the supervisor got wrong. If the student gets most of them right, it is using what it already knows. If it repeats them, it is copying its supervisor.
Step 6 is the part that matters for safety. A future system that simply copied its human supervisors' errors would inherit every blind spot they have.
What the paper found
- Strong students beat weak supervisors. When large models were trained on weak labels, they consistently did better than the models that labeled for them. On language tasks, training GPT-4 on labels from a GPT-2-level model typically recovered about half the gap.
- Naive training is not enough. A large gap remained between the students and the strong ceiling. Results were particularly poor for reward modeling, the task closest to RLHF. The authors conclude that techniques like RLHF may scale poorly to superhuman models without further work.
- Simple methods help. An extra training term that encourages the strong model to be confident, even when it disagrees with its supervisor, raised recovery to nearly 80 percent on language tasks, close to GPT-3.5-level performance. Using intermediate models as stepping stones helped on chess, and extra unsupervised training helped on reward modeling.
The authors are careful about these results. None of their methods worked consistently in all settings, and they call them proofs of concept rather than practical solutions to deploy today.
Where the analogy may not hold
The paper names two important differences between its setup and the real problem.
- Imitation saliency. Future superhuman models may have a strong internal picture of how humans behave, which would make copying human errors easier. Today's strong models may not represent the weak model's errors as well, so avoiding imitation may be easier in the experiment than it will be later.
- Pretraining leakage. Large models learn from human-written text, so human-level skill is already in them, waiting to be drawn out. Superhuman skill was never shown to them in the same way, so it may be harder to draw out.
Where researchers disagree
- The hopeful reading. The authors argue it is feasible to make empirical progress today on a basic challenge of aligning superhuman models. Strong models already generalize past their supervisors, and simple methods improve that.
- The cautious reading. The same results show naive methods leave a large gap, especially for reward modeling. If the two named differences make the real case harder, the experiment may flatter today's methods.
- The open question. Whether a small model's mistakes resemble the mistakes humans will make when supervising a stronger system is not known. The paper itself says the errors weak models make today may differ from human errors.
Learn AI Alignment Theory sets these readings out side by side, with no verdict.
How the courses cover weak-to-strong generalization
Learn AI Alignment Theory covers this in its Advanced course Scalable Oversight, which sits in the level for open research problems and live debates. Its lessons are listed as The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge.

For a related proposal, read our guide to AI safety via debate. The about page shows what each lesson includes.
Frequently asked questions
What is weak-to-strong generalization in simple terms?
It is when a model trained by a weaker supervisor ends up doing better than that supervisor. Researchers use it to study how people might supervise AI that is better than them.
Why use a small model instead of a human?
No superhuman model exists to test with yet. A small model supervising a large one copies the shape of the future problem with systems we have today.
What does the performance gap recovered mean?
It is the share of the gap between the weak supervisor and the strong model's best possible score that the student closes. A score of 1 means the full gap was closed.
Does this mean superhuman AI can be safely supervised?
No. The authors call their methods proofs of concept, note that none worked in every setting, and list ways the real problem may be harder.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, build up through the Basic and Intermediate levels, then take Scalable Oversight for weak-to-strong generalization, debate and amplification, with every side of the debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.