AI safety self-study: a path from the basics to open research

Plenty of people want to understand AI safety, whether to change careers, to work on policy, or simply to follow the arguments properly. Most start with a reading list, open the first paper, and stall somewhere around page six. The material is not too hard. It is that AI safety self-study without structure rarely survives a busy month. This guide gives a concrete order to learn in, the habits that make independent study stick, and what a single study session can look like.

Start with a map, not a reading list

AI safety is several fields sharing one name: technical alignment, evaluation and red teaming, interpretability, control, governance, forecasting and strategy. A reading list hands you all of them at once. A map tells you which ideas the others depend on.

The aisafety.com self-study page collects curricula and reading lists for exactly this kind of independent learning, and its topics are a reasonable map of the field. Whatever you use, settle on one route before you start, so you are not deciding what to read next every evening.

A sensible order to learn in

  1. The core ideas. What alignment means, how modern AI is trained, why objectives are proxies, and how optimisation pressure exposes the gap. None of this needs maths beyond school level.
  2. How training goes right or wrong. Learning from human feedback, why a system can learn a different goal than the one you trained for, keeping systems correctable, and how models are evaluated and red-teamed.
  3. The wider picture. Forecasting AI progress, governance, strategy and economics. Even if you are technical, you need these to judge which problems matter.
  4. Open research. Scalable oversight, interpretability, the science of misbehaviour in language models, agent foundations, and the big unresolved debates.

Go deep in one area of step 4 only after the first three. The open problems make much more sense once you know which simpler ideas they are stuck on.

It also helps to try one hands-on exercise early, so the ideas stop being abstract. Red teaming is a good place to start: if you build or use AI apps, a practical guide to prompt injection testing shows what it looks like to probe a system for behaviour its builders did not intend.

Habits that make AI safety self-study stick

  • Short sessions, often. Twenty minutes most days beats a whole Sunday once a month.
  • Test yourself. A review of study techniques by Dunlosky and colleagues (2013) rated practice testing and spreading study over time as the two most useful methods they examined. Rereading was rated low.
  • Read both sides. Many questions in this field are genuinely open. If you only read one camp, you will learn its conclusions without understanding why others reject them.
  • Go to the source. Summaries are a fine start, but read the original paper for the ideas you will rely on.
  • Keep a glossary. The field has a lot of vocabulary, and much confusion comes from people using the same word differently.

A worked example: one 20-minute session

Habits are easier to keep when you know exactly what a session involves. Here is one, using the habits above.

  1. Minutes 0 to 5: review. Answer the review questions waiting from earlier lessons. Anything you get wrong is a signal, not a failure: it tells you which idea has not stuck yet.
  2. Minutes 5 to 13: one new lesson. Say it is Goodhart's Law. You learn why a measure stops being a good measure once it becomes a target, and you work through an example.
  3. Minutes 13 to 17: say it back. Close everything and write two sentences in your own words: what Goodhart's law says, and one example from your own life, such as a sales target that rewards discounts over profit.
  4. Minutes 17 to 20: pick tomorrow's source. Note one paper from the lesson's sources that you want to open later in the week.

That is all. Repeated most days, it adds up to a solid grounding in a couple of months, with far more retained than a weekend of reading.

Using short lessons alongside the papers

Learn AI Alignment Theory is laid out along that same route. Its 22 courses come in three levels:

  • Basic: What Is Alignment?, How Modern AI Works, Specifying Goals, Agents and Incentives, The Full Risk Landscape.
  • Intermediate: Learning from Humans, Inner Alignment, Corrigibility and Control, Evaluations and Red Teaming, Robustness, Security and Safety Engineering, Forecasting AI Progress, Governing AI, Strategy for the AGI Transition, Economics of Transformative AI, and Ethics, AI Welfare and Good Futures.
  • Advanced: Scalable Oversight, Interpretability, The Science of LLM Misalignment, Agent Foundations, Cooperative AI and Multi-Agent Safety, Research Agendas, Theory and Careers, and The Big Debates.

Every course is open, so you can follow your own route, and the lessons inside a course unlock in order. A topics map shows how the courses line up with the aisafety.com self-study topics.

The habits above are built in. Questions from lessons you finish come back for review at growing gaps of 1, 3, 7, 16, 35 and 90 days, up to five a day, and anything you get wrong returns the next day. Every attempt is graded, and you can retake a lesson to try for a better grade. Debate cards set out each serious position with no verdict, every lesson lists its sources and separates what is known from what is still open, and a glossary covers 349 terms. If a little competition helps you keep going, there is a daily streak and weekly, monthly and all-time leaderboards under a name you choose. The about page has the full picture.

A four-week starting plan

  • Week 1: the first two Basic courses, one lesson a day, plus your daily review.
  • Week 2: the other three Basic courses. Read one source paper from a lesson you found surprising.
  • Week 3: pick the Intermediate course closest to your goal, for example Evaluations and Red Teaming for hands-on work or Governing AI for policy.
  • Week 4: finish it, retake any lesson graded below Good, and choose your next course.

Frequently asked questions

Do I need a technical background for AI safety self-study?

No. The core ideas need no maths beyond school level, and the five Basic courses need no background at all.

How long does it take to get through the material?

The 69 lessons take about eight minutes each, about nine hours in all. Reading the source papers you care about takes longer, and is worth it.

Can I work on AI safety without writing code?

Yes. Governance, forecasting and strategy are large parts of the field, and the Intermediate courses Governing AI and Strategy for the AGI Transition are a good place to start.

Should I learn machine learning first?

Not before you start. The course How Modern AI Works covers what you need for the core ideas, and you can go deeper into machine learning once you know which area you want to work on.

Get started

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.