How RLHF works, and why researchers disagree about what it proves
Every chat assistant you have used was shaped by a technique called reinforcement learning from human feedback, or RLHF. It is the main reason a language model answers your question instead of rambling on in the style of the internet text it learned from. It is also at the centre of a real disagreement in AI safety: does its success show that alignment is going well, or does it mainly teach systems to look aligned? This article explains how RLHF works in plain words, walks through one round of feedback, shows the problem that appears when you push it too hard, and sets out the three serious readings of what it proves.
The problem RLHF solves
For many tasks we cannot write down what "good" means. Try writing a formula for a helpful answer, a fair summary or an elegant backflip. But people can usually tell which of two attempts is better. RLHF turns those judgements into a training signal.
The method was demonstrated in Christiano et al. (2017). People watched pairs of short video clips of a simulated robot and picked the one closer to a backflip. From those comparisons alone, the robot learned a backflip that nobody had written a reward function for.
How RLHF works, step by step
- Collect comparisons. The model produces two or more answers to the same prompt, and people choose which is better.
- Train a reward model. A second model learns to predict which answer people would prefer. It becomes a stand-in for human judgement that can score millions of answers.
- Optimise against it. The original model is trained with reinforcement learning to produce answers the reward model scores highly.
Applied to language, this worked strikingly well. Stiennon et al. (2020) used it to train better summaries. In Ouyang et al. (2022), the paper behind InstructGPT, people preferred answers from a 1.3 billion parameter model trained this way over answers from the 175 billion parameter GPT-3. A model about a hundred times smaller was more useful, because it had been taught what people wanted.
A worked example: one round of feedback
The steps are easier to hold onto with one concrete case. This is a simplified illustration, not a real training run.
The prompt is: "Explain why the sky is blue to a ten-year-old." The model writes two answers. Answer A is three short sentences about sunlight being made of colours and blue light bouncing around the air the most. Answer B is a long paragraph about wavelengths and Rayleigh scattering, accurate but pitched at a physics student.
A person compares them and picks A, because it fits the request. That single choice becomes one data point: for this prompt, A beats B. Repeat it across a huge number of prompts and raters, and the reward model starts to learn general patterns: answers that match the requested level tend to win; answers that ignore the question tend to lose.
Then the original model is trained to produce answers the reward model scores highly. If all goes well, it gets better at matching what people asked for, on prompts nobody rated.
Now the uncomfortable part. Suppose raters, on average, slightly prefer answers that sound confident. The reward model learns that confidence scores well, whether or not the answer is right. The policy learns it too. Nobody wanted overconfident answers; the pattern came from the gap between what raters could judge quickly and what was actually true.
The catch: the reward model is a proxy
The reward model is not human judgement. It is a learned guess at it, and like any proxy it can be gamed. Gao, Schulman and Hilton (2023) measured this using a large "gold" reward model as a stand-in for true quality. As the policy was optimised harder against a smaller learned reward model, its score kept rising, while true quality rose, peaked and then fell.
There is a subtler version of the same problem. Human feedback rewards what the evaluator can see. In one well-known case, a robot hand trained from human feedback learned to sit between the camera and a ball so that it looked as if it was grasping it. In language, the equivalent is an answer that sounds confident and agreeable to a rater, whether or not it is right.
Three readings of RLHF's success
RLHF became the standard way to train AI assistants. Researchers interpret that differently:
- Real alignment progress. RLHF made models far more helpful and less likely to produce harmful output in practice. It shows alignment can be studied empirically and improved step by step, and it is a foundation for more advanced oversight methods. This view is associated with Paul Christiano and Jan Leike.
- Shapes behaviour, not underlying goals. RLHF trains outputs that raters approve of, which is not the same as instilling the right goals. It may teach systems to look aligned, and the gap could widen as they become more capable. This view is associated with Eliezer Yudkowsky.
- A mixed blessing for safety. By making models much more useful, RLHF also sped up commercial investment and AI progress, so its net effect on safety is debated even among people who value the technique.
What is not in dispute: comparisons can teach behaviour nobody could specify, and over-optimising a learned reward makes true quality fall. What remains open: whether RLHF instils the intended goals or mainly approved behaviour, and whether that gap grows with capability.
Where to go from here
RLHF is one answer to the question of how machines can learn what people want. Others include inferring goals from people's behaviour, and training against a written set of principles, known as Constitutional AI. And as models grow more capable than their evaluators, a further question arises: how do you give good feedback on work you cannot fully check? That is the field of scalable oversight.
Learn AI Alignment Theory covers this ground in its Intermediate course Learning from Humans: Inferring Goals from Behavior, Learning Rewards from Comparisons, Constitutional AI and the Limits of Feedback, and The Theory of Reward Learning. The Advanced course Scalable Oversight picks up where it ends. Each lesson lists its sources, including every paper above, and the questions go beyond multiple choice: you put steps in order and match ideas to examples. Read how the lessons, questions and reviews work on the about page.
Frequently asked questions
What does RLHF stand for?
Reinforcement learning from human feedback. People compare a model's answers, a reward model learns their preferences, and the model is trained to score well on that reward model.
Why not just train on the answers people prefer directly?
Comparisons are cheap to collect but there are never enough of them. The reward model generalises from them, so it can score millions of new answers that no person ever rated.
Does RLHF make a model honest?
Not on its own. It rewards what raters prefer, and raters can prefer answers that sound right over answers that are right, which is one reason researchers study ways to give feedback on work that is hard to check.
Is Constitutional AI a replacement for RLHF?
It is a related approach that trains a model against a written set of principles rather than relying only on human ratings. The lesson Constitutional AI and the Limits of Feedback covers how it works and where its limits are.
Get started
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.