How to Read an AI Alignment Paper as a Beginner: A 3-Pass Method

If you want to know how to read an AI alignment paper as a beginner, the short answer is: don't read it front to back on the first try. Read it three times, each time with a different goal. The first pass tells you what the paper is about. The second tells you what it claims and why. The third tests whether you believe it. This post walks you through each pass, with a worked example and a note template you can copy.

Why it is hard to learn how to read an AI alignment paper

Alignment papers mix fields. One paper might combine machine learning experiments, decision theory, and an argument about the future. Authors assume you already know the terms they use, and many terms are new enough that their meaning still shifts between research groups.

There is a second problem. Many papers sit inside a live debate. A paper may be answering an argument you have never seen. Without that context, the paper's main point can look strange or obvious when it is neither.

The fix is not to read harder. It is to read in passes, so you never hold the whole paper in your head at once. The method below adapts the three-pass approach from S. Keshav's short note How to Read a Paper (2007), written for computer science papers, to alignment.

Diagram of the three reading passes, what each one asks and the note it produces, above the three kinds of alignment paper: empirical, theoretical and conceptual

The three kinds of alignment papers: empirical, theoretical, and conceptual

Before pass one, sort the paper into one of three types. The type tells you what kind of evidence to look for.

  • Empirical papers run experiments. They train or test models and report results. Ask: what did they measure, and would the result hold on other models?
  • Theoretical papers use math to prove things. Ask: what are the assumptions, and do they match real systems?
  • Conceptual papers make arguments in plain language. They name a problem, define it, and say why it matters. Ask: is the argument valid, and what would change the author's mind?

Many papers blend types. A conceptual paper may include a small experiment as a demonstration. Still, one type usually carries the main weight. Find it.

Pass 1: Skim the title, abstract, figures and conclusion in 10 minutes

Set a timer for 10 minutes; Keshav suggests five to ten. Read only these parts:

  1. The title and abstract.
  2. The figure captions and the figures themselves.
  3. The section headings.
  4. The conclusion or discussion.

When the timer ends, write one sentence: "This paper claims that ___ because ___." If you can't fill both blanks, that is fine. Write what you have.

Worked example. Say you pick Goal Misgeneralization in Deep Reinforcement Learning (2021). You skim and see its first figure: in the game CoinRun, an agent trained on levels where the coin always sits at the end of the level keeps heading for the end even when the coin is moved, and often skips the coin. Your pass 1 sentence might be: "This paper claims an agent can learn the wrong goal even when its reward was correct, because in training the wrong goal and the right goal looked the same." That one sentence is enough to decide whether to keep going.

Pass 2: Read for the main claim, the setup, and the evidence

Now read the whole paper, but skip proofs and appendices. Keshav gives this pass up to an hour. You are looking for three things.

  • The main claim. Usually in the introduction, often as a list of contributions.
  • The setup. What system, task, or model did they use? What did they hold fixed?
  • The evidence. Which result, proof, or argument supports the claim?

Mark every term you don't know, but don't stop to look each one up. Keep a running list. At the end of the pass, look up the ones that appear more than twice. The rest can wait.

In the goal misgeneralization example, pass 2 shows you the setup: training levels where the coin sat at the end, then test levels where its position was randomized. The evidence is the agent's behavior on the test levels: its skill carried over, its goal did not. Our post on goal misgeneralization covers the idea itself. Now you can ask a sharper question: is this about the agent, or about the training data being too narrow?

Pass 3: Rebuild the argument and test it with your own questions

This is the long pass. Keshav puts it at four or five hours for a beginner, so save it for the papers that matter most to you. Close the paper. On a blank page, rebuild the argument in your own words, step by step. Then open the paper and check where your version differs. Those gaps are where your understanding is thin.

Next, ask four questions:

  1. What would have to be true for this claim to be wrong?
  2. Does the result depend on one model, one task, or one assumption?
  3. Who would disagree, and what would they say?
  4. What does the paper say is still open?

Question three matters most in alignment. Serious researchers disagree about core questions, and a paper rarely states the strongest version of the other side. Try to write that version yourself, then list every assumption the result needs and mark the one you trust least.

This is also how Learn AI Alignment Theory handles debates. Its 71 debate cards set out where researchers disagree, state each serious position fairly, and give no verdict. Each lesson also separates what is known from what is still open, which is the same split you are making in pass 3.

A beginner's jargon cheat sheet: reward hacking, inner alignment, interpretability and more

These terms show up again and again. Short definitions are enough to get through pass 2.

  • Goodhart's Law: when a measure becomes a target, it stops being a good measure.
  • Specification gaming / reward hacking: a system scores well on the goal you wrote down while missing the goal you meant.
  • Side effects: harm a system causes while pursuing its goal, because the goal never said to avoid it.
  • Inner alignment: whether the goal a trained model actually pursues matches the goal it was trained on.
  • Goal misgeneralization: a model keeps its skills in a new setting but pursues the wrong goal there.
  • Deceptive alignment: a hypothesized case where a model behaves well in training in order to pursue a different goal later.
  • Corrigibility: a system's willingness to be corrected or shut down.
  • Interpretability: methods for understanding what is happening inside a model.
  • Scalable oversight: ways for humans to supervise systems that may know more than they do, such as debate or amplification.

Learn AI Alignment Theory has a glossary of 349 terms and a topics map that follows the aisafety.com self-study topics. Both need sign-in. Many of the terms above also have their own lessons. Specifying Goals covers Goodhart's Law, Specification Gaming, Side Effects and Impact, and Tampering and Wireheading. Inner Alignment covers Optimizers Inside Optimizers, Goal Misgeneralization, and Deceptive Alignment and Scheming.

How to take notes and remember what you read (with a one-page paper summary template)

Write one page per paper. No more. Use these headings:

  • Citation: title, authors, year.
  • Type: empirical, theoretical, or conceptual.
  • One-sentence claim: your pass 1 sentence, revised.
  • Setup: what they tested or assumed.
  • Evidence: the key result or argument.
  • Strongest objection: your best version of the other side.
  • Still open: what the paper does not settle.
  • New terms: three at most.
  • Read next: one cited paper you want to follow.

Keep these pages in one searchable place, such as a single notes file with one heading per paper. If you are planning a whole year of reading, our AI safety self-study path suggests an order.

Notes alone fade. What makes ideas stick is being asked about them again, spaced out over time. Learn AI Alignment Theory does this for you: questions from lessons you have finished come back after gaps of 1, 3, 7, 16, 35, and 90 days, up to 5 a day. If you get one wrong, it comes back the next day. You can copy that schedule for your own paper summaries by rereading your "one-sentence claim" line on those days.

Frequently asked questions

Do I need to know machine learning math to read alignment papers?

No, not for most conceptual papers and the main claims of empirical ones. You need the math for theoretical proofs, but you can skip proofs in pass 2 and still follow the argument.

Which AI alignment papers should a beginner read first?

A good place to start is papers that name a problem, such as Concrete Problems in AI Safety (2016) or Risks from Learned Optimization (2019), which introduced the word mesa-optimization. Later papers often assume you know those terms.

Where do alignment researchers publish their work?

Much of it appears on arXiv, the Alignment Forum and AI labs' own research pages, and some at machine learning conferences.

How long should it take to read one alignment paper?

Keshav's guide is about five to ten minutes for pass 1, up to an hour for pass 2, and four or five hours for pass 3 as a beginner. Many papers only deserve pass 1, and that is a good result.

Get started: build your alignment reading habit

Papers are easier when you already know the core ideas they build on. Learn AI Alignment Theory has 22 courses and 69 lessons of about 8 minutes each, across three levels. Basic courses need no background. Intermediate courses cover how today's training methods can go right or wrong. Advanced courses cover open research problems and live debates, including Scalable Oversight, Interpretability, and The Big Debates.

Every lesson lists its sources, so each one points you to papers worth a three-pass read. It is a web app, and you sign in with Google or an emailed code. You can read more on the about page.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.