Agent Foundations Explained: What the Field Studies and Why
This is agent foundations explained. You have probably seen the phrase next to words like "embedded agency" and "decision theory" and wondered how any of it connects to making real AI systems safe. This post gives you a plain answer. You will learn what agent foundations researchers try to understand, why they work with simplified, idealized agents, and how a few classic puzzles show the problems they care about. You will also see where the field touches more practical alignment work.
Agent foundations explained: a plain definition
Agent foundations is the part of AI alignment that asks basic theoretical questions about agents. An agent is anything that has goals, holds beliefs about the world, and picks actions to reach those goals. The field asks what it would mean for such a thing to reason well, choose well, and stay pointed at the goals we gave it.
Most engineering work starts with a system and asks how to fix it. Agent foundations starts one step earlier. It asks what we even mean by "a good reasoner" or "an agent that wants X." The bet is that if those ideas stay vague, we will struggle to check whether a powerful system has them.
So the outputs are usually not code. They are definitions, toy models, thought experiments, and sometimes proofs.
Why alignment researchers study idealized agents
Physicists study frictionless planes. Nobody thinks real planes lack friction. The simple case lets you see the core law before the mess gets added back.
Agent foundations does the same with agents. Researchers imagine an agent with unlimited compute, or a perfect model of its environment, and ask how it should behave. Then they look for places where even that ideal agent breaks. If a perfect reasoner gets confused by some setup, a messy trained neural network probably will not handle it gracefully either.
Embedded agency: when the agent is part of the world it models
Abram Demski and Scott Garrabrant's Embedded Agency starts from a gap in the classic picture. Traditional models treat the agent as cleanly separated from its environment, acting on it from outside, able to model it in every detail, and never needing to reason about itself.
Real agents are not like that. An AI runs on hardware inside the world it models. Demski and Garrabrant survey what follows: such an agent must rely on models that fit inside the environment they model, and must reason about itself as just another physical system, made of parts that can be changed and can work at cross purposes. The embedded agency post goes through it in full; here is the short version:
Embedded agency raises problems the clean model hides:
- The world is bigger than the agent. An agent cannot hold a full model of a world that contains itself, because the model would have to contain a model of the model, and so on.
- The agent can change itself. If it can edit its own goals or reward signal, what does it mean for it to "want" anything stable?
- The agent is made of parts. Those parts might pursue their own aims.
The second point links to reward tampering and wireheading: an agent that can reach its own reward channel may prefer to tamper with it. In the clean console model that option does not exist.

Decision theory: how an agent should choose (CDT, EDT, and FDT)
Decision theory asks a simple question: given what you know, which action should you take? The hard part is defining what "the result of an action" means. Three families of answer come up most often.
- Causal decision theory (CDT): pick the action that causes the best outcome.
- Evidential decision theory (EDT): pick the action that, once you learn you took it, is the best news about the outcome.
- Functional decision theory (FDT): proposed by Eliezer Yudkowsky and Nate Soares in a 2017 paper, it treats your decision as the output of a fixed mathematical function and asks which output of that function would give the best outcome.
A worked example: Newcomb's problem
Here is the classic case. Work through it yourself before reading the answers.
- A predictor that has been right nearly every time sets out two boxes.
- Box A is clear and holds $1,000.
- Box B is opaque. It holds $1,000,000 if the predictor guessed you would take only Box B. It is empty if the predictor guessed you would take both.
- The prediction is already made. You now choose: Box B alone, or both boxes.
A CDT agent reasons: the money is already in place, my choice now cannot change it, so taking both boxes always adds $1,000. It takes both. And because the predictor is good, it usually walks away with $1,000.
An EDT agent reasons: people who take only Box B almost always find a million in it. It takes one box. FDT also takes one box, but for a different reason: the predictor ran a model of your decision procedure, so choosing "one box" is choosing what that model output.
Yudkowsky and Soares argue FDT gets more utility than CDT on this problem; defenders of CDT and EDT dispute that framing.
Why does this matter for AI? Software agents can be copied, simulated, and predicted far more easily than humans. An AI may face other agents that have read its code. Which decision theory it uses changes how it bargains, cooperates, and responds to threats. There is no consensus here, and serious researchers defend each position.
Logical uncertainty and reasoning about yourself
Ordinary probability handles uncertainty about facts in the world, like whether it will rain. It assumes you know all the logical consequences of what you believe. Real reasoners do not. You might not know the millionth digit of pi, even though it is fixed. That is logical uncertainty.
For AI this shows up whenever a bounded system has to hold sensible beliefs about math and code it has not finished checking. Scott Garrabrant and colleagues' Logical Induction gives one answer: a computable algorithm that assigns probabilities to every statement in a formal language and refines them over time, learns patterns in logical truths before it can check them, and holds accurate beliefs about its own beliefs while avoiding the standard paradoxes of self-reference.
A related question is self-trust. If an agent builds a successor, how can it be confident the successor reasons correctly, when proving that may be as hard as proving it about itself?
Robust delegation, corrigibility, and subsystem alignment
The last cluster is about passing goals from one agent to another without losing them.
Robust delegation asks how a principal (you) can hand a task to a stronger agent and trust the result. The agent knows more than you, so you cannot just check every step. This is the same gap that motivates process supervision versus outcome supervision.
Corrigibility asks how to build an agent that accepts correction or shutdown, even though most goals give an agent a reason to avoid being switched off. Writing down a utility function that is indifferent in the right way, and not exploitable, turns out to be surprisingly hard.
Subsystem alignment asks what happens when an optimizer creates another optimizer inside it. Training searches for a model that scores well (the mesa-optimization post covers it), and the model it finds might itself be pursuing some internal goal. That inner goal may match training only by accident. This is the theoretical root of worries about hidden behavior like sleeper agents.
How agent foundations connects to practical alignment work like reward misspecification and value alignment
Agent foundations can look abstract, but many practical problems are its questions in disguise.
- Reward misspecification is what happens when the goal you wrote down differs from the goal you meant. Agent foundations asks how to define goals so an ideal optimizer would not exploit that gap.
- Value alignment needs some account of what an agent "values" in the first place. That is a foundations question.
- Inner alignment problems, such as goal misgeneralization and deceptive alignment, come straight from the subsystem alignment picture above.
Frequently asked questions
Who works on agent foundations?
The work cited here comes from a fairly small set of authors: Demski and Garrabrant on embedded agency, Yudkowsky and Soares on functional decision theory, and Garrabrant and colleagues on logical induction. The questions themselves are not owned by any one group.
Is agent foundations still relevant in the age of large language models?
This is an open debate. Some researchers argue the theory predicts how capable systems will behave whatever their architecture; others argue effort is better spent studying the trained models we actually have.
How is agent foundations different from interpretability research?
Interpretability studies the insides of real trained networks to see what they compute. Agent foundations studies idealized agents to clarify concepts like goals, beliefs, and choice, which interpretability can then look for.
Get started
Learn AI Alignment Theory has an Advanced course called Agent Foundations, alongside related courses on Inner Alignment, Corrigibility and Control, and The Big Debates. Lessons run about 8 minutes, and each one separates what is known from what is still open, lists its sources, and includes hands-on activities where you switch assumptions on and off to see what changes.
Where researchers disagree, such as on decision theory or the relevance of idealized agents, debate cards lay out each serious position fairly with no verdict. Questions from finished lessons come back on a spaced schedule, so ideas like embedded agency stick. If you are new, start with the Basic level, which needs no background. You sign in with Google or an emailed code. You can read more on the about page.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.