Scalable oversight: how humans could supervise smarter AI
Today, much of AI training depends on people judging whether an answer is good. That works while people can tell. Scalable oversight is the research problem of keeping that check working when the AI is better at the task than the person checking it.
This guide explains where the problem comes from, the main ideas researchers have proposed, and why they disagree about how far those ideas will go.
The checking problem
Methods like reinforcement learning from human feedback ask people to compare answers and pick the better one. Our guide to how RLHF works explains the details. The method quietly assumes the person can spot the better answer.
That assumption breaks in two ways. Some tasks are simply too expensive to check often: Concrete Problems in AI Safety (Amodei and colleagues, 2016) listed "scalable supervision" as one of five practical safety problems for exactly this reason. And some answers may be too technical or subtle for any one person to judge, however much time they have.
What scalable oversight means
Measuring Progress on Scalable Oversight for Large Language Models by Samuel Bowman and colleagues (2022) gives a clear definition: the problem of supervising systems that potentially outperform us on most skills relevant to the task at hand.

The same paper names the practical difficulty. We do not yet have systems that broadly exceed our abilities, so it is hard to test oversight methods on the very situation they are meant for.
Their answer is to study tasks where human specialists succeed but both unaided non-specialists and current AI systems fail. The specialists supply a trusted answer key. Everyone else plays the part of a future overseer facing a system that knows more than they do, and researchers can measure whether a given method closes the gap.
The main proposals
Most proposals share one idea: give the human overseer help, often from AI systems themselves.
- AI assistance. A person checks an answer while talking with a model. Bowman and colleagues found that on two question-answering tests, people chatting with an unreliable AI assistant did substantially better than both the model alone and the people on their own.
- Debate. In AI safety via debate (Irving, Christiano and Amodei, 2018), two AI systems take turns making short arguments, and a person judges which side gave the most true, useful information.
- Amplification. Iterated Amplification (Christiano and colleagues, 2018) builds up a training signal for a hard problem by combining answers to easier pieces of it.
- Reward modeling. Jan Leike and colleagues (2018) outlined a research direction of learning a reward function from interaction with the user and training on it, and discussed the challenges of scaling it to complex tasks.
A worked example: running an oversight experiment by hand
Bowman and colleagues describe an experimental design you can follow on a small scale. The quiz and roles below are made up for illustration.
- Pick a task with three levels of skill. Take a set of hard questions about a long legal contract. A contract lawyer gets them right. A non-lawyer reading quickly does not, and neither does the AI model on its own.
- Record the baselines. Score the non-lawyer alone, and the model alone. Say each gets about half right.
- Give the non-lawyer an assistant. They may question the model as much as they like, but the model is sometimes wrong and sometimes confidently wrong.
- Score the pair. If the non-lawyer with the assistant beats both baselines, the oversight method added something: the person found the right answers without being a lawyer.
- Check against the expert. The lawyer's answers are used only to score, never to help. That is what makes the setup a stand-in for overseeing a system smarter than us.
The useful part is step 5. Because a real expert exists, you can measure whether a method lets a non-expert reach the truth, which is the thing scalable oversight is meant to do.
What the early results show
The early results are encouraging but limited. Bowman and colleagues call their finding an encouraging sign that the problem can be studied with present models.
A related line of work asks a different version of the question. Weak-to-Strong Generalization (Burns and colleagues, 2023) used a weaker model to supervise a much stronger one, as a stand-in for humans supervising superhuman systems. The strong models did better than their weak supervisors. With an extra training trick, a GPT-4 model supervised by a GPT-2-level model came close to GPT-3.5-level performance on language tasks. The same paper says they were still far from recovering the strong models' full abilities.
Where researchers disagree
- The tractable view. The authors of these papers see a problem that can be worked on now, with experiments like the one above, and expect steady progress from better methods.
- The worried view. Burns and colleagues note that techniques like RLHF may scale poorly to superhuman models without further work. If an overseer cannot tell good answers from convincing ones, every method built on human judgment inherits the gap.
- The open questions. Irving and colleagues say plainly that whether debate works depends on empirical questions about humans and the tasks we want AI to do. The same holds for the other proposals: whether toy and stand-in experiments predict the real case is not yet known.
Learn AI Alignment Theory sets out these positions side by side, with no verdict.
How the courses cover scalable oversight
Learn AI Alignment Theory has a full Scalable Oversight course in its Advanced level, which covers open research problems and live debates. Its lessons are The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge.

If you are planning a route through the field, our AI safety self-study path shows where oversight fits. The about page lists all 22 courses.
Frequently asked questions
What is scalable oversight in simple terms?
Finding ways for people to check and correct AI systems even on tasks where the AI is better than they are.
Why can't people just check the answers?
Some tasks are too costly to check often, and some answers may be too technical or subtle for a person to judge on their own.
Is debate the same as scalable oversight?
No. Debate is one proposed method for scalable oversight, alongside AI assistance, amplification and reward modeling.
Has scalable oversight been solved?
No. Early experiments are encouraging, but the authors themselves describe them as first steps on stand-in tasks.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, build up through the Basic and Intermediate levels, then take the Scalable Oversight course, with every side of the debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.