AI-Written Critiques That Help Human Evaluators, Explained

Training AI with human feedback only works if people can tell good answers from bad ones. As answers get longer and more technical, that gets harder. AI-written critiques are one proposed fix: a model reads an answer and writes down what is wrong with it, and a person uses that critique to judge. This post covers what the two main studies found, how critique-assisted review works, its limits, and how it fits into the bigger oversight picture.

What AI-written critiques are

A critique, in this research, is a short piece of plain text pointing out problems in an answer: a missing fact, a wrong claim, a bug in code. The critic is a language model trained to write such comments. The human evaluator still makes the call; the critique is help, not a verdict.

The aim is to raise the quality of human feedback, so the models trained on that feedback learn the right lessons. McAleese and colleagues put the problem directly in 2024: reinforcement learning from human feedback "is fundamentally limited by the capacity of humans to correctly evaluate model output."

The oversight problem behind AI-written critiques

This is a piece of scalable oversight, the problem of supervising systems that may outperform us on most skills relevant to the task. If a person cannot spot a subtle bug, they may rate buggy code as fine, and training then rewards code that looks right rather than code that is right. See how RLHF works for why ratings matter so much.

Critiques try to narrow that gap without asking the person to become an expert. Finding a flaw someone else has pointed to is easier than finding it unaided.

What the research shows

Self-critiquing models (2022)

Saunders and colleagues fine-tuned models to write critiques of summaries. Their critiques "help humans find flaws in summaries that they would have otherwise missed." That held for flaws that occurred naturally and for flaws planted on purpose in summaries written by people trying to mislead.

Two more findings matter. Larger models wrote more helpful critiques, and on most tasks they were better at critiquing their own work. Larger models could also use their own critiques to improve their summaries. But the authors add a caution: "even large models may still have relevant knowledge they cannot or do not articulate as critiques." A critic can know something is wrong and still not say it.

CriticGPT (2024)

McAleese and colleagues trained critic models on code from real assistant tasks. On code with naturally occurring model errors:

  • Model-written critiques were preferred over human critiques in 63% of cases.
  • The models caught more bugs than human contractors paid to review code.
  • The critics found hundreds of errors in ChatGPT training data that had been rated "flawless," even though most of those tasks were not code.
Diagram of critique-assisted evaluation: a model answer, a critic model's comments, and a human evaluator making the final rating, with the main findings

A worked example: critique-assisted review

Here is how the workflow runs, step by step, on a coding answer. The details are a made-up illustration of the setup the papers describe.

  1. The task. A user asks an assistant for a function that removes duplicates from a list while keeping the original order.
  2. The answer. The assistant returns code that looks clean and runs.
  3. The critique. A critic model writes: "This converts the list to a set, which does not keep the original order."
  4. The check. The evaluator tests the claim on a short list such as 3, 1, 3, 2. If the order changes, the critique is right.
  5. The rating. The evaluator rates the answer down and notes the bug. That rating becomes training data.

Step 4 is the important habit. A critique is a claim, and claims can be wrong, which leads to the limits.

Limits: hallucinated bugs and teams

The CriticGPT paper is plain about its weak spot: critics can produce "hallucinated bugs that could mislead humans into making mistakes they might have otherwise avoided." A confident but false critique can turn a correct answer into a rejected one.

Its answer is teams. Human contractors working with critics caught similar numbers of bugs to the critics alone, "while hallucinating less than LLMs alone." The person catches the critic's inventions; the critic catches what the person misses.

A second limit is the critic itself. If the model being checked and the model doing the checking share blind spots, critiques can miss the same problems. Baker and colleagues (2025) show a related risk in monitoring: when reasoning models were trained hard against a monitor, they "learn obfuscated reward hacking," hiding their intent while still cheating. Any checker that becomes a training target can be gamed.

How critiques fit with debate and weak judges

Critiques share a family resemblance with other oversight ideas. In AI safety via debate, two models argue and a judge picks. With critiques, one model points out flaws and a person checks them. Both try to make a weaker judge more accurate, which is what weak judge experiments measure.

There is an encouraging sign for the weaker-checks-stronger case: Baker and colleagues found that GPT-4o, a weaker model, could effectively monitor the reasoning of o3-mini, a stronger one, for reward hacking. Researchers differ on how far this carries. One reading is that AI assistance can keep human feedback useful as models improve. Another stresses the open limits these papers name themselves: knowledge a critic does not voice, hallucinated bugs, and checks that get gamed under pressure.

Frequently asked questions

Do AI-written critiques make human evaluators more accurate?

In the studies covered here, yes for the tasks tested: critiques helped people find flaws they would otherwise miss, and human and critic teams caught as many bugs as critics alone with fewer false alarms.

Can a critic be wrong?

Yes. The CriticGPT paper reports hallucinated bugs that can mislead people, which is why a person checks each claim.

How is this different from AI debate?

Debate has two models argue opposite sides for a judge. Critique assistance has one model point out problems in an answer, and a person decides.

Get started: learn scalable oversight step by step

On Learn AI Alignment Theory, the Advanced course Scalable Oversight has lessons on The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge. The Intermediate course Learning from Humans covers Constitutional AI and the Limits of Feedback. Lessons run about 8 minutes, every lesson lists its sources and separates what is known from what is still open, and debate cards set out each serious position with no verdict. You sign in with Google or an emailed code.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.