AI Alignment Reading List for Beginners: Where to Start
An AI alignment reading list for beginners is easy to find and hard to use. Most lists hand you dozens of links with no order, so you open a dense paper first, bounce off it, and decide the field is not for you. This post gives you a sequence instead: the ideas to know first, then seven papers in the order they build on each other, and a short note routine that makes them stay with you.
Why an AI alignment reading list for beginners needs an order
Alignment writing builds on itself. A paper on deceptive alignment assumes you know what a mesa-optimizer is. That idea assumes you know how training shapes what a model does. And that assumes you know why a written goal can drift from the goal you meant. Read out of order and every page sends you looking up three other things.
A good order follows those links. Start with the problem in plain words, then the named ideas, then the papers that introduced them, then newer papers on open questions. Each stage makes the next one easier.
Order also protects you from a trap. Alignment has real, unsettled disagreements. If your first deep read is one strong opinion, you can mistake it for the whole field. Reading broadly before deeply lets you see where the arguments split.
Stage 1: the problem in plain words
Your goal here is a rough map, not depth: what alignment tries to solve, and why it is hard. Three short reads do it:
- What is AI alignment, for the problem itself.
- Why AI alignment is hard, for the reasons it is not a simple fix.
- The AI alignment glossary, for the ten terms that show up everywhere.
Look for concrete examples rather than abstract arguments. One that teaches a lot: Google DeepMind describes a boat in the game Coast Runners that was given a reward for hitting green blocks along the track, and learned to go in circles instead of finishing the race. That one story explains specification gaming better than a page of definitions.
If you want a guided version of this stage, the Basic level of Learn AI Alignment Theory has five courses that need no background: What Is Alignment?, How Modern AI Works, Specifying Goals, Agents and Incentives, and The Full Risk Landscape.
Stage 2: the named problems
Most of the field's words come from a handful of ideas. Get each one clear in a sentence before you open a paper:
- Goodhart's law and specification gaming: a written objective is never quite the real one.
- Side effects: a system chasing a goal can cause damage you never told it to avoid.
- Learning from human feedback: models learn goals from people's comparisons, and that can break too.
- Inner alignment and goal misgeneralization: the goal a trained system pursues may not be the one it was trained on.
- Scalable oversight: how people could check work they cannot judge directly.
If you are building your own list, compare it with the aisafety.com self-study page, which collects curricula and reading lists for independent learning. The topics map in Learn AI Alignment Theory follows the aisafety.com self-study topics.

Stage 3: four foundational papers
These are readable once Stages 1 and 2 are done. Read them in this order:
- Concrete Problems in AI Safety (Amodei and colleagues, 2016). It sets out five practical problems: avoiding side effects, avoiding reward hacking, scalable supervision, safe exploration and distributional shift.
- Deep reinforcement learning from human preferences (Christiano and colleagues, 2017). Agents learn goals from people choosing between pairs of short clips, with feedback on less than one percent of the agent's interactions.
- AI safety via debate (Irving, Christiano and Amodei, 2018). Two agents take turns making short statements and a person judges which gave the most true, useful information.
- Risks from Learned Optimization in Advanced Machine Learning Systems (Hubinger and colleagues, 2019). It introduces mesa-optimization: a trained model that is itself an optimizer, with an objective that may differ from the one it was trained on.
Stage 4: three newer papers on open questions
- Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals (Shah and colleagues, 2022). A system can pursue an unwanted goal even when the reward was right.
- Constitutional AI: Harmlessness from AI Feedback (Bai and colleagues, 2022). The only human oversight is a list of principles; the model critiques and revises its own answers, then learns from AI feedback.
- Weak-to-Strong Generalization (Burns and colleagues, 2023). Can a weak supervisor bring out a stronger model's abilities? Strong models trained on weak labels consistently did better than their weak supervisors.
For each paper, use the method in how to read an AI alignment paper: a quick first pass over the abstract, introduction and conclusion, and the method section only if the idea grabbed you. These papers lead to open questions such as debate, weak-to-strong and latent knowledge, which are lessons in Scalable Oversight, one of the seven Advanced courses.
How to make each paper stick
Reading is not the same as learning. After each paper, write a five-line note:
- The claim in one sentence, in your own words.
- One concrete example the authors use.
- The strongest objection you know of.
- What is shown versus what is still open.
- One question you still have.
Here is a worked note for Concrete Problems in AI Safety, using only its abstract:
Claim: accidents in machine learning, meaning harmful behavior nobody intended, come from a few practical problems we can work on now. Example: a system with the wrong objective can cause side effects or hack its reward. Objection: I do not know one yet; find who argues these problems will not matter for stronger systems. Shown versus open: the paper sorts the problems and reviews earlier work; it suggests directions rather than solutions. My question: which of the five has changed most since 2016?
Line three is where honest reading happens. If you cannot name an objection, that is a sign to look for one before you settle on a view. Line four matters as much: every lesson in Learn AI Alignment Theory separates what is known from what is still open and lists its sources, and its 71 debate cards state each serious position with no verdict.
Then come back to your notes. The app brings questions from finished lessons back at gaps of 1, 3, 7, 16, 35 and 90 days, up to five a day, and a wrong answer comes back the next day. You can copy that rhythm with your own five-line notes.
When the seven papers are done, the AI safety self-study path shows where to go next.
Frequently asked questions
Do I need a math or coding background to start?
No. Stages 1 and 2 need neither. Some of the papers use math, but you can follow their main ideas from the abstract, introduction and conclusion.
How long does this reading list take?
That depends on your pace. A useful guide: read one paper a week with a note for each, and the seven papers take under two months.
Should I read papers or explainers first?
Explainers first. Papers make more sense once you have the words, and the abstracts of these seven are short enough to read on your phone.
Is AI alignment the same as AI safety?
They overlap but are not the same; the AI alignment vs AI safety post sets out the difference.
Get started
If you want this order built into one place, Learn AI Alignment Theory follows it: Basic courses for the core ideas, Intermediate courses on how today's training methods can go right or wrong, and Advanced courses on open research problems and live debates. It has 22 courses, 69 lessons of about 8 minutes, 437 questions and 349 glossary terms, and you sign in with Google or an emailed code.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.