AI Alignment Research Agendas Compared: A Clear Guide for 2026

Search for ai alignment research agendas compared and you mostly find long reading lists or essays arguing for one camp. This guide is narrower. For five major agendas it says what failure each one is trying to prevent, what it has to assume, the paper that sets it out, and where its hard open problem sits. By the end you should be able to place a new paper in the right bucket and ask it the right skeptical question.

What a research agenda is, and three questions to ask

A research agenda is a bet: this part of the problem matters most, and this method can make progress on it. Agendas differ less in their goals than in their assumptions, and that changes what counts as progress. A result that looks like a big step to one group can look beside the point to another.

Ask three questions of any agenda:

  • Which failure is it trying to prevent? Goals we cannot state, goals hidden inside a model, work we cannot check, or a model working against us.
  • What does it assume? About future systems, about people, about what can be measured.
  • How would you know it worked? A theory, a measurement, a passed test.
Diagram of five AI alignment research agendas side by side: agent foundations, interpretability, scalable oversight, learning from feedback and AI control, each with the failure it targets and the bet it makes

Five agendas, five different failures, five different bets.

Agent foundations: theory before trust

Agent foundations asks what good reasoning even means for an agent that lives inside the world it is reasoning about. Abram Demski and Scott Garrabrant's "Embedded Agency" (2019) starts from the point that standard models of rational action treat the agent as cleanly separated from its environment. A real agent is not: it must use models that fit inside the world it models, and reason about itself as a physical system made of parts. More in agent foundations explained.

The bet: we may build systems we do not understand at a conceptual level, so a clear theory is needed first. The open problem: critics say the abstractions are hard to connect to real neural networks; supporters say that gap is the reason the work matters.

Interpretability: reading what models compute

Interpretability tries to open the model and find the mechanism behind its behavior, instead of judging it only by outputs. The difficulty is set out in Anthropic's "Toy Models of Superposition" (2022): networks often pack many unrelated concepts into a single neuron, which makes them harder to read, and the paper builds a toy model where this can be fully understood. See mechanistic interpretability for beginners.

The bet: large networks contain structure people can understand, and understanding it will expose problems that behavior alone hides. The open problem: whether today's methods scale to the largest models, and whether partial understanding is enough.

Scalable oversight: AI helping people check AI

Scalable oversight starts from a worry: how do you check work you cannot judge yourself? Geoffrey Irving, Paul Christiano and Dario Amodei's "AI safety via debate" (2018) notes that human judging can fail when a task is too complicated, and proposes two agents taking turns making short statements while a human judges which gave the most true, useful information. "Weak-to-Strong Generalization" (Burns and colleagues, 2023) studies an analogy: strong models trained on a weak model's labels consistently did better than their weak supervisors. More in scalable oversight explained.

The bet: AI can help people evaluate AI. The open problem: the debate paper itself says whether it works depends on empirical questions about humans and tasks, such as whether an honest argument wins against a persuasive false one.

Learning from feedback: RLHF and constitutions

This agenda is closest to how today's chat models are trained: learn what people want from their comparisons or from written principles, and train toward it. In "Constitutional AI" (Bai and colleagues, 2022), the only human oversight is a list of rules or principles; a model critiques and revises its own answers, then is trained with reinforcement learning from AI feedback. More in how RLHF works.

The bet: feedback, from people or from principles, captures what we want. The open problem: the Weak-to-Strong paper points out that RLHF relies on humans being able to supervise model behavior, and future models may be too complex for that. A model can also score well on feedback and still learn a different goal.

AI control: safety even if alignment fails

AI control takes a different angle. Ryan Greenblatt, Buck Shlegeris and colleagues' "AI Control" (2023) notes that researchers had not tested whether safety techniques still hold if the model is deliberately trying to subvert them. They build and test "protocols" using a powerful untrusted model (GPT-4), a weaker trusted one (GPT-3.5) and a little trusted human labor. More in AI control vs AI alignment.

The bet: we can contain or catch a misaligned model. The open problem: whether control keeps working as the untrusted model becomes much more capable than the trusted one.

AI alignment research agendas compared: a worked example

Take one claim, such as "interpretability will catch deceptive models before they act". Then, on paper:

  1. Write what interpretability needs to be true for the claim to hold (readable structure, methods that scale).
  2. Write what a scalable oversight researcher might say (a debate partner could expose the deception instead).
  3. Write what an AI control researcher might say (assume it is not caught, and limit what the model can do).
  4. Now switch off one assumption at a time, such as "models have readable structure", and see which agendas still support the claim.

Doing this a few times trains you to see an agenda as a set of assumptions, not a team to join. Notice that the agendas mostly fail in different places. Some researchers take that as a reason to work on several at once; others argue that focus matters more. That disagreement is open.

Frequently asked questions

Which AI alignment research agenda is most promising?

Researchers disagree, and the answer depends on what you expect future systems to look like. List each agenda's assumptions, decide which you find most plausible, then read the strongest case against your view.

Do I need math for alignment theory?

Agent foundations leans heavily on math and decision theory; interpretability, evaluations and control lean more on coding and experiments. Pick the agenda that matches the skills you want to build.

How do theory and experiments help each other?

Theory can say what to look for, and experiments can show whether a theory describes real systems. Each also catches the other's blind spots.

Are these the only agendas?

No. Evaluations and red teaming, governance, and cooperative AI are agendas too; these five are a starting map, not a complete one.

Get started

Learn AI Alignment Theory follows these agendas through its courses. The Advanced level has Agent Foundations, Interpretability, Scalable Oversight (with lessons on The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge), Research Agendas, Theory and Careers, and The Big Debates. The Intermediate course Learning from Humans includes Constitutional AI and the Limits of Feedback, and Corrigibility and Control is its own course. Its 71 debate cards set out where researchers disagree, each position stated fairly with no verdict, and some activities let you switch assumptions on and off. Read more on the about page, then start learning by signing in with Google or an emailed code.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.