RLAIF Explained: Training AI with AI Feedback, Step by Step
RLAIF stands for reinforcement learning from AI feedback. This post gives you RLAIF explained step by step: it works like the familiar RLHF pipeline, except that a language model, not a person, decides which of two answers is better. You will see why the method exists, how each step works, what a direct comparison with RLHF found, and which alignment questions it leaves open.
What is RLAIF? Training AI with AI feedback in plain terms
Every feedback-based training method has to answer one question: who decides what a good answer looks like? In RLHF, people compare pairs of model outputs and pick the better one. In RLAIF, a separate model makes that comparison, working from written instructions about what "better" means.
The rest of the pipeline stays much the same. You collect many judgments, train a reward model to predict them, and then use reinforcement learning to push the main model toward answers the reward model scores highly. Only the source of the labels changes.
Why RLAIF exists: the cost of human labels
The 2023 paper "RLAIF vs. RLHF" by Harrison Lee and colleagues at Google DeepMind and Google opens with the problem: RLHF works, but gathering high-quality human preference labels is expensive. A model can label far more pairs than a team of people can.
The idea did not start there. The paper credits RLAIF to Bai and colleagues, whose Constitutional AI paper (2022) begins: "As AI systems become more capable, we would like to enlist their help to supervise other AIs." In that work, the only human oversight of harmlessness was a written list of rules or principles, and the authors say the methods need far fewer human labels.
The open question is whether AI labels are good enough, and whose judgment they really carry. That links to a wider debate covered in AI value alignment: whose values count and why.
RLAIF explained step by step
Here is the pipeline from the Lee paper, with an example you can follow.

Canonical RLAIF and direct RLAIF, after Lee and colleagues (2023).
Step 1: Generate two answers
Take a set of prompts. For each one, sample two answers from the model you want to improve. Say the prompt is "Summarize this 600-word news article in three sentences." You get Summary 1 and Summary 2.
Step 2: Ask an AI labeler to compare them
A labeler model gets instructions, the article and both summaries, and is asked which is better. A simple labeling prompt might read:
A good summary is accurate, covers the main points, and adds nothing the article does not say. Here is the article, then Summary 1 and Summary 2. Consider the accuracy and coverage of each summary and explain which one is better. Then answer: Preferred summary = 1 or 2.
Three details from the paper matter here:
- Position bias. The order in which the answers are shown can sway the labeler, especially smaller ones. So each pair is labeled twice, with the order reversed the second time, and the two results are averaged.
- Soft labels. Instead of a hard pick, the paper reads how likely the labeler was to say "1" or "2", giving a preference such as 0.6 and 0.4.
- Reasoning first. Asking the labeler to explain before it answers generally made its labels agree more with human labels; detailed instructions and worked examples helped only on some tasks.
Step 3: Train a reward model
Repeat across thousands of prompts, then train a reward model on the AI preferences: it reads a prompt and an answer and outputs a score. Learning a reward from pairwise comparisons is its own subject, with its own theory.
Step 4: Fine-tune with reinforcement learning
Optimize the main model to produce answers the reward model rates highly. The paper describes an optional KL penalty, which keeps the model close to where it started and guards against "reward hacking", where the model finds answers that score well without being good.
The direct variant
The paper also introduces direct RLAIF (d-RLAIF). During RL, the labeler scores each answer from 1 to 10, and that score is the reward. This skips the reward model. The authors' reason: a reward model trained on the starting model's answers can go "stale" as the trained model's answers move away from that data.
RLAIF vs RLHF vs Constitutional AI
- RLHF: people compare outputs, a reward model learns from those comparisons, and RL optimizes against it.
- RLAIF: the same pipeline, but a model makes the comparisons.
- Constitutional AI: a method with two phases. First, the model critiques and revises its own answers, and it is fine-tuned on the revisions. Then a model compares pairs of answers, a preference model learns from those AI preferences, and RL uses it as the reward. The paper names that second phase "RL from AI Feedback".
Put simply, RLAIF is the general technique, and Constitutional AI is one recipe that uses it. In both, people do not leave the loop: they write the principles or the labeling prompt, choose the labeler, and judge the result.
What the research found: where AI feedback matched human labels
Lee and colleagues tested RLAIF against RLHF on summarization, helpful dialogue and harmless dialogue, with people judging the final models.
- Against the starting supervised model, people preferred RLAIF 71 percent of the time and RLHF 73 percent of the time on summarization, and 63 and 64 percent on helpful dialogue. The paper calls this comparable.
- On harmless dialogue, RLAIF had a harmless rate of 88 percent, RLHF 76 percent and the starting model 64 percent.
- RLAIF still beat the starting model when the labeler was the same size as the model being trained, or even the same checkpoint.
- Direct RLAIF did better than the standard version.
Read these results with care. They cover specific tasks where a capable model can judge quality fairly well. On another dataset the authors tried, neither RLHF nor RLAIF showed meaningful gains once they corrected for longer answers. And "comparable to RLHF" means the AI labels did about as well as human labels, not that either was enough.
Open questions and alignment risks of RLAIF
These are concerns researchers raise about training on any learned judge, not findings of the Lee paper.
- The labeler's quirks become the model's. Whatever the labeler prefers gets trained in, blind spots included. Position bias is one quirk the paper measured and corrected; others may go unnoticed.
- Goodhart's law still applies. The trained model optimizes a proxy. Our post on reward model overoptimization sets out what happens when you push a learned reward too hard.
- Errors line up. Different people make different mistakes. One labeler model can make the same mistake across a whole dataset.
- Only the output is judged. The labeler sees the answer, not why the model gave it. Hidden behavior of the kind covered in sleeper agents in AI is exactly what output-based feedback can miss.
One view, in the Constitutional AI paper's own words: "As AI systems become more capable, we would like to enlist their help to supervise other AIs." The question the other way is how far you can trust a labeler that is itself a trained model. Both deserve a fair hearing.
Frequently asked questions
Is RLAIF the same as Constitutional AI?
No. RLAIF is the general idea of using AI preference labels for reinforcement learning. Constitutional AI is a specific method that uses RLAIF in its second phase, guided by written principles, after a self-critique and revision phase.
Can RLAIF replace human feedback?
Not fully in the studies above: people still wrote the principles or prompts and judged the final models. The Lee paper found AI labels comparable on three tasks, which does not cover tasks where judging quality is itself hard.
Does direct RLAIF skip the reward model?
Yes. The labeler scores each answer from 1 to 10 during training and that score is the reward. This avoids a stale reward model, but the labeler has to run throughout training.
Why does RLAIF matter for scalable oversight?
Scalable oversight asks how to supervise AI systems on tasks people cannot easily judge. RLAIF is an early, concrete case of AI helping to supervise AI, so its strengths and failures are a guide for methods like debate and amplification.
Get started
On Learn AI Alignment Theory, the Intermediate course Learning from Humans has lessons called Learning Rewards from Comparisons, Constitutional AI and the Limits of Feedback, and The Theory of Reward Learning. The Basic course Specifying Goals has Goodhart's Law and Specification Gaming, and the Advanced course Scalable Oversight has The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge.
Each lesson takes about 8 minutes, lists its sources and separates what is known from what is still open. Debate cards set out each serious position with no verdict, and hands-on scenarios let you switch assumptions on and off. You sign in with Google or an emailed code. Read more on the about page.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.