Process Supervision vs Outcome Supervision Explained
If you have read about how AI models are trained to reason, you have probably met two terms that sound alike: process supervision and outcome supervision. This is process supervision vs outcome supervision explained in plain terms: how each one works, what the two main studies found, why alignment researchers care, and where the open questions sit. One worked example runs through the whole post so you can see the difference, not just read about it.
Process supervision vs outcome supervision explained in one example
Both are ways to give a model feedback while it learns. They differ in what gets graded.
- Outcome supervision grades only the final result. Right answer, reward. Wrong answer, none.
- Process supervision grades each step of the reasoning, so a bad step is marked bad even inside a solution that ends well.
Here is the example. You ask a model to simplify the fraction 16/64. It writes:
- Start with 16/64.
- Cancel the 6 on top with the 6 on the bottom.
- That leaves 1/4.
The answer, 1/4, is correct: 16 times 4 is 64. The method is nonsense. You cannot cancel digits like that; it only works here by luck. Try the same move on 12/24 and you get 1/4, when the answer is 1/2. An outcome grader sees a right answer and gives full marks. A process grader sees that step 2 is invalid and marks it wrong. That gap is the whole debate in one small problem.

How outcome supervision works
Outcome supervision is the simpler method. You need problems and a way to check the final answer. For math, the check can be automatic: compare the model's number with the known solution. For code, you can run tests.
- Cheap labels. You only need the right answer, not an expert reading every line.
- Easy to scale. Automatic checks can grade a great many attempts.
- Any path allowed. The model is free to find whatever route works.
The cost is a thin signal. One right-or-wrong at the end says nothing about which step helped or hurt. And it rewards any path to the right answer, including broken ones like the digit cancelling trick. If a shortcut reaches the target, outcome supervision reinforces the shortcut.
How process supervision works
Process supervision needs someone to judge each step. Because people cannot read every solution a model writes, the step labels are usually used to train a separate model, a process reward model, which then scores new reasoning step by step.
- The model writes step-by-step solutions to many problems.
- People read the steps and mark each one as good or bad.
- A process reward model learns to predict those marks.
- That model is then used to pick the best solutions, or as the reward during training.
In the fraction example, a process reward model trained on good labels would flag step 2, so the whole solution scores low even though the answer is right. The feedback points at the exact place the reasoning went wrong.
What the studies found on math reasoning
Jonathan Uesato and colleagues ran a comparison on GSM8K, a set of grade school math word problems, in 2022. Their abstract reports that pure outcome supervision reached similar final-answer error rates with less labeling. But to get correct reasoning steps, they needed process supervision, or a learned reward model that imitates it. Their best results cut final-answer errors from 16.8% to 12.7%, and reasoning errors among solutions with a correct final answer from 14.0% to 3.4%. That second number is the 16/64 problem: right answer, wrong path.
Hunter Lightman and colleagues followed with Let's Verify Step by Step in 2023, on the harder MATH dataset of competition problems. They report that process supervision significantly outperformed outcome supervision there, and their process-supervised model solved 78% of problems in a representative subset of the MATH test set. They also found that active learning, choosing which solutions people label, significantly improved process supervision. And they released PRM800K, the full set of 800,000 step-level human labels behind their best reward model, so others can study the method.
Keep the scope in mind. Both studies are on math, where a step is fairly easy to check. Whether the results carry over to messier tasks is an open question.
Why process supervision matters for alignment
Fewer rewarded shortcuts
When you reward only outcomes, you invite the model to game the measure: a model graded on "the tests pass" might learn to change the tests. That is reward misspecification at work. Grading steps narrows the room for shortcuts, because the shortcut itself gets flagged.
Reasoning people can follow
Process supervision rewards steps that people can read and agree with. If those visible steps really drive the answer, reading them tells you what the model is doing. That links it to chain of thought monitoring, which asks whether we can read AI reasoning at all.
The cost of safety
An "alignment tax" is the performance you give up to make a system safer. If safety always costs ability, there is pressure to skip it. In these math studies, the method that checks the reasoning was also as accurate or more accurate on final answers, so here the tax was not a cost. Whether that holds as tasks get harder is a separate question, and one study area says little about others.
Limits and open problems
Process supervision is not a finished answer. Here are the main open issues, with the case on each side.
- Labeling cost. Step labels need skilled people reading long solutions. Supporters point to active learning and to reward models that, once trained, label on their own. Skeptics note the hardest tasks are where expert time is scarcest.
- Steps that are not the real reason. A model's written steps may not be why it reached its answer. If so, grading the steps grades a story, and rewarding good-looking steps could train more convincing stories. How often this happens is an active research question.
- Steps people cannot judge. Process supervision assumes a person can tell a good step from a bad one. For tasks beyond that, it runs into the oversight gap that scalable oversight tries to close.
- Gaming the grader. A process reward model is a learned proxy too. Gao and colleagues showed that pushing too hard against a learned reward model can hurt the true score; the reward model overoptimization post explains it. Grading steps moves that risk rather than removing it.
Both kinds of supervision sit inside the larger story of training from people's judgments, which the how RLHF works post covers.
Frequently asked questions
Is process supervision the same as chain of thought prompting?
No. Chain of thought prompting asks a model to show its work when it answers. Process supervision is a training method that grades that work step by step, so it relies on written steps but is a separate idea.
Does process supervision prevent reward hacking?
It reduces some forms, because an invalid shortcut gets flagged at the step where it appears. It does not remove the problem: a process reward model is itself a proxy that a strong optimizer could learn to fool.
What is a process reward model?
A model trained on people's step-by-step labels to score each step of a solution. Once trained, it can rank candidate solutions or act as a reward without a person reading every line.
When is outcome supervision the better choice?
When the final result is easy to check automatically and the path matters less, or when step labels cost too much. Uesato and colleagues found it reached similar final-answer error rates with less labeling.
Get started
Process and outcome supervision sit where several bigger topics meet: gaming a reward, learning from people's feedback, and overseeing systems that may outgrow their graders. Learn AI Alignment Theory covers those pieces in lessons of about 8 minutes.
- The Intermediate course Learning from Humans has Learning Rewards from Comparisons, Constitutional AI and the Limits of Feedback, and The Theory of Reward Learning.
- The Advanced course Scalable Oversight has The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge.
Every lesson lists its sources and separates what is known from what is still open. Debate cards set out where researchers disagree, each position stated fairly with no verdict, and hands-on activities let you switch assumptions on and off. You sign in with Google or an emailed code. Read more on the about page.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.