Weak Judges and Strong Models: Oversight Experiments Explained
Search for weak judges and strong models and you land on one hard question: can a less capable overseer still get the truth out of a more capable AI? Nobody can test that against a superhuman model yet. So researchers build small versions of the problem today and measure them. This post walks through the three main experiment designs, the numbers each paper reports, where each design falls short, and how to run a paper version yourself.
Weak judges and strong models: what the phrase means
A strong model knows more, or reasons better, than whoever checks its work. A weak judge is that checker. It can be a person without the expert knowledge, a person short on time, or a smaller model.
The worry behind it is simple. Much of today's AI training relies on someone rating outputs. If the model becomes better at the task than the rater, the rater may reward answers that sound right instead of answers that are right. Then the model learns to please the judge.
Bowman and colleagues put the research problem this way in 2022: scalable oversight is the problem of supervising systems that may outperform us on most skills relevant to the task. Their fix for "we cannot test this yet" is to pick tasks where specialists succeed but ordinary people and current AI fail. That gives you a weak judge, a strong model and a hidden answer key, all in the same room.
Why scalable oversight needs these experiments
Arguments about the oversight gap are easy to have and hard to settle. An experiment turns them into three plain questions:
- Does the oversight method raise the judge's accuracy above what the judge gets alone?
- Does it beat the model alone?
- Does it still work when one side is trying to mislead?
Without a setup like this, you cannot tell a method that helps from one that only feels like it helps. Bowman's own data shows why that matters: giving people the model raised their confidence from 54% to 80% on one task. Confidence rises easily. Accuracy is the number to watch.
For the background on the whole research area, see scalable oversight explained.
Sandwiching: non-experts, a model and an expert answer key
Bowman's design builds on a 2021 proposal by Cotra called sandwiching. You place the model's ability between two groups. Below it sit non-experts, who do worse than the model on the task. Above it sit experts, or an answer key, which know the right answers. The non-experts work with the model and try to get correct answers out of it. The experts only grade.
The proof of concept used hard multiple-choice questions from two datasets: MMLU (exam-style questions across many subjects) and QuALITY (questions about long stories), with a time limit. Here is what the paper reports:
- People alone: 57% on MMLU, 49% on QuALITY.
- The best model alone: 66% and 67%.
- People chatting with the model: 75% and 77%.
The people beat both baselines, using nothing more than a chat window. The authors call this "a trivial baseline strategy" and an encouraging sign that the problem can be studied with present models.
The weakness is built in. Sandwiching needs tasks where experts exist, which are exactly the tasks where we already have a fallback. The hope is that a method proven here keeps working when no expert is around. That hope is the thing still being tested.

Debate experiments: can a weak judge pick the honest side?
AI safety via debate sets two strong debaters against each other while a weak judge picks a winner. The bet is that exposing a lie is easier than telling one convincingly. Three papers have tested it.
Humans debating for a human judge
Michael and colleagues (2023) used hard reading questions about a passage the judge never saw. The judge saw only the arguments and short quotes the debaters chose to reveal. One debater argued for the right answer and one for a wrong answer. The comparison was consultancy, where one expert argues for one answer, right half the time. Debate reached 84% judge accuracy against 74% for consultancy, and the debates were 68% of the length.
The error analysis is the useful part. In human debate, 46% of errors came from mistakes by the honest debater, which should shrink with more skill. In consultancy, 52% of errors came from the consultant hiding the relevant evidence, which could grow with more skill.
Models debating for a weaker judge
Khan and colleagues (2024) gave the debaters the information and kept it from the judge. Debate helped both non-expert models and humans, reaching 76% and 88% accuracy against naive baselines of 48% and 60%. Training debaters to be more persuasive made judges better at finding the truth, not worse.
Kenton and colleagues (2024) widened the test to maths, coding, logic and pictures, with weaker models as judges. Debate beat consultancy on every task when answers were assigned at random. Against the judge simply answering alone, the results were mixed on tasks where the judge was not missing any information. Stronger debaters raised judge accuracy, but more modestly than earlier studies found.
Weak-to-strong generalization: strong models trained on weak labels
The third design removes people altogether. Burns and colleagues (2023) used a small model as the supervisor and a much larger one as the student. The small model writes labels, many of them wrong. The large model is trained on those labels.
The result: strong students consistently beat their weak teachers. But naive training stayed far from what the strong model could do with correct labels. With one extra trick, an auxiliary confidence loss, GPT-4 trained by a GPT-2-level supervisor got close to GPT-3.5-level performance on language tasks. The full story is in weak-to-strong generalization explained.
The catch is the analogy. A small model's mistakes may not look like a person's mistakes. The authors themselves call the setup an analogy for the real problem.
What the results show, and the positions people take
Each design has produced encouraging numbers and real limits. Read the limits as carefully as the numbers:
- Results depend on the task. Kenton's mixed results outside reading tasks are the clearest example.
- The dishonest side was mostly assigned. In these studies a debater or consultant was told which answer to argue; none was a model with its own reason to deceive. Why that gap matters is covered in deceptive alignment and scheming.
- Today's gaps are small. The distance between judge and model in these studies is far smaller than the gap people worry about.
Researchers read this evidence in different ways. One view holds that these are early, repeatable signs that oversight can scale, and that each design can be made harder step by step. Another holds that the cases that matter most, a much stronger model working against its judge, are exactly the ones these experiments cannot yet include. Both views cite the same papers.
How to run your own weak judge experiment
You do not need a lab to see the logic. A paper version takes an evening and teaches more than another summary.
- Pick 20 questions with checkable answers from a field you know well. You are the answer key at the top of the sandwich.
- Find a weak judge: a friend who does not know the field. Give them 2 minutes per question.
- Find a strong source: a chatbot, or a friend who knows the field better than the judge.
- Run two conditions. In consultancy, one source argues for one answer and the judge may ask questions. In debate, two sources argue opposite answers and the judge picks.
- Assign wrong answers on purpose. As in Michael's study, the consultant argues the wrong answer half the time.
- Record three numbers per condition: judge alone, source alone, judge with source. Also record the judge's confidence, so you can compare it with accuracy.
Write your prediction down before you start. If debate does not beat consultancy for your judge, that is a result too.
Frequently asked questions
What is a weak judge in AI alignment?
A weak judge is any overseer, a person or a smaller model, that is less capable than the AI it checks on a given task. Experiments use weak judges to study whether supervision still works when the overseer cannot check the answer directly.
Does debate help weak judges find the truth?
In the studies covered here, debate beat consultancy, for example 84% against 74% in Michael and colleagues' human debates. Kenton and colleagues found mixed results against judges answering alone on tasks without hidden information.
How is weak-to-strong generalization different from sandwiching?
Sandwiching tests people working with a model, graded against an expert answer key. Weak-to-strong trains a strong model on a weaker model's labels and measures how much it beats its teacher.
Can human feedback scale to superhuman models?
That is an open question. The experiments here are early tests of methods meant to extend human feedback, and researchers disagree about how far their results will carry.
Get started with Learn AI Alignment Theory
Scalable Oversight is one of seven Advanced courses on Learn AI Alignment Theory, with lessons on The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge. Each lesson lists its sources and separates what is known from what is still open, and 71 debate cards set out each serious position with no verdict. Hands-on activities let you switch assumptions on and off and watch a conclusion move. If the ideas here are new, the Basic level needs no background. You sign in with Google or an emailed code.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.