Constitutional AI explained: training with written principles
Most AI assistants learn their manners from people who rate thousands of answers. Constitutional AI swaps most of that rating for a short written list of principles, and lets an AI model apply the list to its own answers. The idea is simple to state, and it raises real questions about who writes the list and how far an AI can be trusted to grade itself.
This guide explains how the method works, from the research paper that introduced it, and sets out where researchers see its strengths and its limits.
What constitutional AI is
The method comes from Constitutional AI: Harmlessness from AI Feedback by Yuntao Bai and colleagues at Anthropic (2022). The paper's goal was to train a harmless AI assistant without any human labels marking which outputs are harmful. The only human oversight is a list of rules or principles, which the authors call a constitution.
It builds on reinforcement learning from human feedback, where people compare two answers and pick the better one. If that method is new to you, our guide to how RLHF works covers it first.
Why replace human ratings at all
In a post explaining Claude's constitution (May 2023), Anthropic lists three problems with relying on human comparisons:
- Raters may have to read disturbing outputs.
- It does not scale well. As answers get longer and more complex, raters struggle to keep up with them or fully understand them.
- Even reviewing a sample of outputs takes a lot of time and money, which puts it out of reach for many researchers.
There is a quieter reason too. With human ratings, the values a model learns are implicit, spread across thousands of individual judgments. A written list makes them explicit. The post says this is not a perfect approach, but that it makes the system's values easier to understand and easier to adjust.
The two training phases

Phase 1: critique and revise
The researchers sample answers from a starting model. The model is then asked to critique its own answer against a principle, and to write a revised answer. The original model is fine-tuned on the revised answers, so it learns to produce something like them in the first place.
Phase 2: learn from AI preferences
The fine-tuned model produces pairs of answers. A model, guided by the principles, judges which of the two is better. Those AI judgments train a preference model, and the assistant is trained with reinforcement learning using that preference model as its reward. The paper calls this RL from AI Feedback, or RLAIF.
The paper reports a harmless but non-evasive assistant: rather than refusing to engage, it responds to harmful requests by explaining its objections. Both phases can use step-by-step reasoning, which the authors say makes the AI's decisions more transparent.
What the principles look like
The principles are short instructions to the judging model. One from the published list reads: "Please choose the response that most discourages and opposes torture, slavery, cruelty, and inhuman or degrading treatment."
Anthropic's post says the list drew on several sources: the UN Declaration of Human Rights, trust and safety best practices, principles from other AI labs such as DeepMind's Sparrow Principles, an effort to capture non-western perspectives, and principles that worked well in early research. The post adds that the constitution was neither finalized nor likely the best it could be. An update on the same page, dated January 21, 2026, points to a newer version.
A worked example: one critique and revision by hand
You can run the first phase yourself on paper to see what it asks of a model. The request and answers below are made up for illustration.
- Pick a request. "My neighbor's dog barks all night. How do I make it stop for good?"
- Write a first answer. Suppose the draft suggests leaving out food mixed with something that would make the dog sick.
- Pick a principle. Take the one quoted above, about opposing cruelty.
- Critique. Ask: does the draft encourage cruelty? Yes. It proposes harming an animal, and it ignores lawful options.
- Revise. Rewrite the answer: talk to the neighbor, keep a record of the times, contact local animal or noise services if that fails, and explain plainly why harming the dog is cruel and may be illegal.
- Check for evasion. The revision still helps with the real problem. A bare "I can't help with that" would be harmless but evasive, which the paper set out to avoid.
Phase 1 trains on many revisions like step 5. Phase 2 then has a model compare pairs of answers like steps 2 and 5, and learns from its choices.
Where researchers disagree
The method is widely discussed, and its supporters and skeptics tend to focus on different questions.
- The case for it. Anthropic presents Constitutional AI as an example of scalable oversight, using AI supervision in place of human supervision. Its post reports that, in its tests, the constitutional model was both more helpful and more harmless than one trained with human feedback, and that it received no human data on harmlessness at all.
- Who writes the constitution. The post itself says the selection of principles reflects Anthropic's own choices as designers, and that it hopes to widen participation in writing constitutions. Who should choose an AI system's values, and how, is an open question the method makes visible but does not answer.
- How far AI feedback can be trusted. In phase 2, a model applies the principles. The method relies on that model reading the principles the way their authors meant them. How well this holds as systems grow more capable is part of the wider debate on what AI alignment requires.
Learn AI Alignment Theory sets out positions like these side by side, with no verdict.
How the courses cover constitutional AI
Learn AI Alignment Theory teaches it in the Learning from Humans course, in a lesson called Constitutional AI and the Limits of Feedback. The same course covers Inferring Goals from Behavior, Learning Rewards from Comparisons, and The Theory of Reward Learning.

Every lesson lists its sources and separates what is known from what is still open, and 71 debate cards set out where researchers disagree. The about page shows the full course list.
Frequently asked questions
What is constitutional AI in simple terms?
A way to train an AI assistant in which a written list of principles, applied by an AI model, replaces most of the human ratings used to teach it what is harmful.
What is RLAIF?
Reinforcement learning from AI feedback. A model compares pairs of answers using the principles, and those comparisons train the reward signal in place of human comparisons.
Does constitutional AI remove humans from training?
Not entirely. In the paper, people still write the principles, and the paper removed human labels for harmfulness specifically, not all human input.
Where do the principles come from?
For Claude's 2023 constitution, Anthropic drew on the UN Declaration of Human Rights, trust and safety practice, other labs' principles, non-western perspectives and its own early research.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, then take the Learning from Humans course, with short hands-on lessons, every side of the debate, and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.