Recursive Reward Modeling Explained Simply, With Examples

Here is recursive reward modeling explained simply: you train an AI by learning what a person wants from their feedback, and when the task gets too hard for the person to judge alone, you give them AI helpers that were trained the same way on easier tasks. The method builds on itself one level at a time. This post covers the idea, a worked example, how it relates to other oversight proposals, and the open questions its own authors name.

What is recursive reward modeling?

Start with two pieces. A reward model learns to score behavior the way a person would. The agent is the AI that does the task, trained to get high scores from that reward model. Jan Leike and colleagues, in their 2018 paper Scalable agent alignment via reward modeling, call this pair reward modeling: learn a reward function from interaction with the user, then optimize it with reinforcement learning.

The same paper proposes the recursive step. Agents trained with reward modeling on simpler tasks, in narrower areas, help the user evaluate the work of a more capable agent in a more general area. The paper lists the kinds of help: giving relevant extra information, summarizing large amounts of data, interpreting the agent's internals, and solving sub-problems the user has carved off. With that help, the user gives feedback that trains the next agent, which can in turn help with the level above.

The problem it solves: tasks too hard to check

Most training from human feedback assumes a person can look at an answer and tell whether it is good. That works for "is this email polite?" It breaks for "is this large codebase free of security holes?"

That is the oversight gap: what AI systems can do may grow faster than our ability to check it. If you reward what looks good to a rushed person, you can train a system that is good at looking good. The scalable oversight post covers the gap and the family of answers to it. Recursive reward modeling is one answer to a narrow question: how do you keep a person's judgment in charge when the work outgrows it?

The bet underneath: evaluating is easier than doing

Leike and colleagues state the key assumption plainly: for many tasks, evaluating an outcome is easier than producing the right behavior. You may not be able to write a correct proof, but if a helper points to the step that looks wrong, you can often check that one step.

They also name the exception. For a yes or no question, checking the answer is no easier than giving it. Their fix is to ask for an explanation as well, since judging an explanation is usually easier than producing one.

Diagram of recursive reward modeling: a helper agent from the level below assists the user, whose feedback trains the reward model for the next agent

Recursive reward modeling explained simply, level by level

Here is the recursion as a short procedure you can follow:

  1. Level 1. Train an agent with reward modeling on a task a person can judge directly, such as "does this function match its description?"
  2. Level 2. The task is bigger: "review this whole module." The person cannot check it alone, but the Level 1 agent flags mismatches and summarizes what the code does. With that, the person compares two reviews and picks the better one. Those choices train a Level 2 reward model and a Level 2 reviewer.
  3. Level 3. The task is "review this whole system." The person now works with Level 2 reviewers as helpers, and their feedback trains Level 3.

Notice what the person never has to do: check every line personally. They judge pieces they can actually check, with helpers pointing to the right pieces. The paper adds that in practice it may make more sense to train all the agents together rather than strictly one after another.

A worked example you can try on paper

Pick a task you could not fully check in an hour, such as a 40-page technical report. Now list the helpers that would make your judgment trustworthy:

  • A summarizer that gives you one page in plain words.
  • A checker that re-does each calculation and flags any that do not follow.
  • A critic that writes the strongest objection to each main claim.

For each helper, ask: could I judge its output directly? If yes, it can be trained with plain reward modeling, and it sits one level down. If no, it needs helpers of its own. That is the recursion.

The critic is not just a thought experiment. William Saunders and colleagues trained models to write critiques of summaries, and the critiques helped people find flaws they would otherwise have missed, including flaws planted to mislead. Larger models wrote more helpful critiques. They call this a proof of concept for AI-assisted human feedback, not the full recursive scheme.

How it compares to amplification, debate and RLHF

RLHF is the base loop with no recursion: people compare outputs, a reward model learns from them, and the agent is trained against it. The how RLHF works post covers it.

Iterated amplification, from Paul Christiano, Buck Shlegeris and Dario Amodei, builds up a training signal for hard problems by combining solutions to easier sub-problems. Leike and colleagues say recursive reward modeling can be thought of as an instance of it, using reward modeling where amplification used supervised or imitation learning. The iterated amplification post walks through that one.

AI safety via debate takes another route: two AIs make short statements and a person judges which gave the most true, useful information. Debate relies on two AIs competing; recursive reward modeling relies on helpers cooperating with the judge. The AI safety via debate post explains it.

Where it can fail: the open questions

The authors are direct that this is a research direction, not a plan that achieves alignment. They name the questions:

  • Do errors pile up? Do the mistakes of a narrower helper lead to larger mistakes in training the next agent, or can training be set up so small mistakes shrink, for example with groups of agents that check each other?
  • Loopholes in the reward model. If the reward model captures most of the goal but not all of it, the agent may find bad solutions. Recursion adds more reward models that could each be slightly off; see reward misspecification.
  • Cost and drift. Is the needed feedback affordable, and does the reward model hold up when the situations change?
  • Unacceptable outcomes. How do you stop a harmful action before it happens, rather than learning from it after?

Saunders and colleagues add a caution of their own: even large models may know things they cannot or do not put into critiques. A helper that keeps quiet is a gap the person cannot see. Researchers disagree about how far these problems can be solved, and each side has serious arguments.

Frequently asked questions

Who proposed recursive reward modeling?

Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini and Shane Legg, in their 2018 paper on scalable agent alignment via reward modeling.

Is recursive reward modeling the same as RLHF?

No. RLHF trains a reward model from direct human feedback. Recursive reward modeling uses the same loop but gives the person AI helpers, trained the same way, so they can judge tasks too hard to judge alone.

Has anyone tested parts of it?

Pieces of it, yes. Saunders and colleagues showed model-written critiques helped people find flaws in summaries; the full multi-level scheme is still a proposal.

Does it solve AI alignment?

Its own authors say it is not a plan that achieves alignment. It is a proposal for keeping human oversight useful as tasks get harder, with open questions about errors piling up and loopholes.

Get started

The Advanced course Scalable Oversight in Learn AI Alignment Theory covers the surrounding ideas: The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge. For the base loop, the Intermediate course Learning from Humans has Learning Rewards from Comparisons, Constitutional AI and the Limits of Feedback, and The Theory of Reward Learning.

Lessons take about 8 minutes. Every lesson lists its sources and separates what is known from what is still open, debate cards set out each serious position with no verdict, and hands-on activities let you switch assumptions on and off. You sign in with Google or an emailed code. Read more on the about page.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.