Math for AI Alignment: What You Need to Start
If you are looking into math for AI alignment, the short answer is: less than you fear, and more than zero. Four areas cover most of what you will meet: probability, linear algebra, calculus with optimization, and decision and game theory. You do not need all of them on day one. You can start on alignment ideas now and pick up each piece of math when an idea calls for it. This guide shows what each area does in alignment work, with a small worked example for each, and a three-month study order.
Math for AI alignment: what you need, and why it matters
Alignment is about making AI systems do what we intend. Many of its core problems can be stated in plain words, for example "a system optimizes a measurable stand-in instead of the thing you care about." But to tell how bad a problem is, or whether a proposed fix works, you need numbers. Math turns a worry into a claim you can check.
How much you need depends on the work you want to do:
- To follow the arguments: school algebra, a feel for probability, and the idea of a function you push uphill or downhill.
- To read empirical papers: working linear algebra, basic statistics and gradient descent.
- To do theory: comfort with proofs, decision theory and game theory.
Most people start in the first group, and that is fine. The Basic level of Learn AI Alignment Theory has five courses that need no background, so you can build the math alongside them.

Four areas of math, in the order of the study plan below.
Probability: reading test results honestly
Alignment work keeps asking "how sure are we?" Will this model behave badly in rare cases? Did the safety test measure what we think? Probability gives you the tools.
Worked example. You run 100 independent tests of a model and it fails none. Is its failure rate zero? Not necessarily. Suppose the true rate were 3 percent. The chance of passing one test is 0.97, so the chance of passing all 100 is 0.97 to the power 100, about 0.048, or roughly 5 percent. Unlikely, but not rare enough to rule out. At a true rate of 1 percent, a clean run of 100 happens about 37 percent of the time. So a clean result tells you failures are uncommon, not that they never happen, and a model used millions of times could still fail often in absolute numbers.
Evaluation has a further catch: a model might act differently when it can tell it is being tested. Our post on evaluation awareness covers that. Your numbers are only as good as the assumption that the test matches real use.
What to learn: conditional probability, Bayes' rule, expected value, distributions and confidence intervals. The "estimate a number" questions in the lessons are good practice: you commit to a guess, then see how far off you were.
Linear algebra: the language of neural networks
A neural network does a lot of matrix multiplication. A layer takes a vector of numbers, multiplies it by a matrix of weights and passes the result on. Interpretability research asks what those vectors mean, so linear algebra is the first thing you need to read that work.
The key ideas are vectors, dot products, matrices and projections, and the idea to really own is a direction. The "linear representation hypothesis", set out by Park and colleagues in 2023, is the idea that high-level concepts are represented as directions inside a model; their paper connects it to linear probing and to steering, and finds such directions in LLaMA-2.
Worked example. Say a model's internal state is the vector h = (2, 1, 0), and you suspect the concept "French" is the direction f = (1, 0, 0). The dot product h · f = 2×1 + 1×0 + 0×0 = 2 measures how far h points along f. A linear probe learns a direction like f and compares it with many internal states. If the scores separate French text from English text, the direction carries that information.
What to learn: vectors and dot products, matrix multiplication, basis and dimension, and eigenvectors at a basic level. Then mechanistic interpretability is a good next topic.
Calculus and optimization: how models learn from feedback
Models learn by gradient descent. You define a loss, a number that says how wrong the model is, then nudge every weight in the direction that lowers it. That direction comes from derivatives, so you need the idea of a derivative and the chain rule. You do not need hard integrals.
Optimization is also where goals go wrong. The optimizer pushes on whatever you wrote down, not on what you meant. This is Goodhart's law, a lesson in the Basic course Specifying Goals: when a measure becomes a target, it stops being a good measure.
Worked example. In Christiano and colleagues' 2017 paper on learning from human preferences, a person sees two clips and picks the better one. The model predicts the chance clip 1 is preferred as exp(r1) / (exp(r1) + exp(r2)), where r is a learned reward. The paper says this follows the Bradley-Terry model. It is the same as sigmoid(r1 - r2), with sigmoid(x) = 1 / (1 + e^(-x)). If r1 = 2 and r2 = 1, then sigmoid(1) is about 0.73: a 73 percent chance people prefer clip 1. Training adjusts the rewards so these predictions match real choices. See how RLHF works for the next step.
The Intermediate course Learning from Humans has lessons called Learning Rewards from Comparisons and The Theory of Reward Learning, and Inner Alignment has one called Optimizers Inside Optimizers.
What to learn: derivatives, partial derivatives, the chain rule, gradient descent and the sigmoid function.
Decision theory and game theory: the core of alignment theory
Decision theory asks how an agent should choose, given its beliefs and preferences. Game theory asks what happens when several agents choose at once.
Worked example. Hadfield-Menell and colleagues' "The Off-Switch Game" (2016) studies a human who can press a robot's off switch and a robot that can disable it. Make it concrete: the robot expects 10 units of value from its task, and 0 if switched off. If disabling the switch costs nothing, a robot that takes its goal for granted does at least as well by disabling it. The paper shows such agents have an incentive to disable the switch, except when the human is perfectly rational, and that for the robot to want to keep its off switch, it needs to be uncertain about the value of the outcome and to treat the human's choice as evidence. That is the seed of corrigibility: how do you build an agent that does not resist correction?
Game theory matters once there are many AI systems, or AI systems and people together. The prisoner's dilemma shows how two agents each acting sensibly alone can reach an outcome both dislike. Among the courses, the Basic course Agents and Incentives and the Advanced courses Agent Foundations and Cooperative AI and Multi-Agent Safety are the ones whose titles point at agents and their choices.
What to learn: expected utility, payoff tables, Nash equilibrium, and the idea of one agent modeling another. Some logic and proof reading helps for agent foundations.
A study order for your first three months
This plan pairs each piece of math with alignment ideas, so the math has a reason to exist. It fits beside an AI safety self-study path.
- Month 1: foundations. Probability and expected value, plus vectors and dot products. Alongside them, the five Basic courses: What Is Alignment?, How Modern AI Works, Specifying Goals, Agents and Incentives, and The Full Risk Landscape.
- Month 2: how models learn. Derivatives, the chain rule, gradient descent and the sigmoid. Pair them with the Intermediate courses Learning from Humans, Inner Alignment, and Evaluations and Red Teaming.
- Month 3: agents and theory. Decision theory and basic game theory. Pair them with Corrigibility and Control, then try an Advanced course such as Interpretability or Agent Foundations.
Keep sessions short: a lesson takes about 8 minutes, so one fits around a math chapter. Questions from finished lessons come back after gaps of 1, 3, 7, 16, 35 and 90 days, up to 5 a day, and a wrong answer returns the next day. When a term trips you up, look it up in the 349-term glossary.
Frequently asked questions
Do I need a math degree to work on AI alignment?
The ideas in this guide can be learned on your own. How deep you go depends on the work: theory needs more proof skill than evaluation or policy work. Our post on alignment research careers sets out the paths.
What math does alignment theory use beyond machine learning?
Decision theory, game theory and some logic. They show up in work on corrigibility, agent foundations and cooperation between AI systems.
Can I learn alignment ideas before the math?
Yes. Ideas like Goodhart's law and specification gaming make sense in plain words first, and the math then sharpens them. Hands-on activities with sliders let you change an assumption and see the result before you write an equation.
Which area should I learn first?
Probability is a good first pick, because it helps you read results, from evaluations to training curves. Linear algebra comes next if interpretability draws you in.
Get started
Learn AI Alignment Theory runs from first principles to open research, in Basic, Intermediate and Advanced levels, with a topics map of the aisafety.com self-study topics the courses follow. Each lesson lists its sources and separates what is known from what is still open. You sign in with Google or an emailed code.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.