Embedded agency explained: an AI inside its own world
Embedded agency is the problem of how an agent should reason and act when it is part of the world it is trying to change. Most theories of rational choice quietly assume the agent sits outside its world, like a player holding a game controller. A real AI system runs on hardware inside the world, can be changed by it, and can change itself. Researchers think that gap hides several unsolved puzzles.
This guide explains the idea with the paper's own example, walks through its four open problems, gives you a short exercise, and sets out why researchers disagree about how much it matters.
What embedded agency means
The idea is set out in Embedded Agency by Abram Demski and Scott Garrabrant, posted to arXiv in 2019 and also available as a sequence on the AI Alignment Forum. They call the usual picture of an agent "dualistic": the agent is made of different stuff from its environment and touches it only through fixed channels.
An embedded agent breaks that picture in four ways. In the paper's words, it does not have well-defined input and output channels, it is smaller than its environment, it can reason about itself and self-improve, and it is made of parts similar to the environment. The authors stress that these are tangled together, not separate.
Alexei and Emmy: the picture behind the idea
The paper opens with two characters.
- Alexei is playing a video game. The game has clear inputs and outputs. Alexei can hold the whole game in his head. He never has to think about himself, because he is not inside the game. He can treat himself as an unchanging atom.
- Emmy is playing real life. She is inside the world she is trying to improve. Her choice is just one more fact about that world. She is smaller than the world, so she cannot model all of it. She is built from the same parts as everything else, so she can study herself, change herself, and break herself.
The paper sums up the difference in one line: "Alexei can poke the universe and see what happens. Emmy is the universe poking itself."
The formal model of Alexei is AIXI, Marcus Hutter's theoretical agent. Under some assumptions it does reasonably well in every computable environment, but it is itself uncomputable, so it could never fit inside the worlds it reasons about. That is exactly the assumption an embedded agent cannot make.

The four open problems of embedded agency
Demski and Garrabrant split the puzzle into four subproblems.
- Decision theory. Choosing normally means comparing what happens under each option. But if you are part of the world, and could in principle prove which action you will take, what does it mean to ask what would happen if you did something else? The paper calls this the problem of logical counterfactuals. It also covers worlds that contain copies of the agent or accurate predictions of it.
- Embedded world-models. Standard Bayesian reasoning starts with a set of possible worlds that includes the true one, then rules out the wrong ones. An agent smaller than the world cannot even hold one complete, correct model of it. The paper lists logical uncertainty and ontological crises, which is what to do when you find your goals were written in the wrong concepts.
- Robust delegation. An agent that builds a smarter successor, or improves itself, wants the result to keep its goals. The successor faces the reverse problem: how to learn the goals of a weaker, inconsistent creator. Subproblems include value learning and corrigibility. This chapter also covers Goodhart's law and wireheading.
- Subsystem alignment. A powerful search can create new optimizers by accident, and these may pursue a slightly different goal while looking aligned. The paper connects this to mesa-optimizers, the learned optimizers described by Hubinger and colleagues.
The authors say these are not four separate problems. They are different faces of one confusion about what agency is.
A worked example: four questions for a real system
You can try this on paper in ten minutes. Pick an AI system you know. Here we use an imagined coding assistant that runs on a developer's laptop, can read and edit files, and can start other programs. Ask the paper's four questions.
- Does it have clean input and output channels? Not really. It reads files that it may have written earlier, and the programs it starts change the same machine it runs on. Its actions loop back into its own inputs.
- Is it smaller than the world it models? Yes. It cannot hold the whole codebase, the internet, and the people using it in memory at once. It works from partial models, so it must reason well without a complete picture.
- Can it reason about and change itself? If its settings or instructions are files on the same laptop, it can in principle edit them. Now ask what should stop it changing its own goals, and whether you would want it to accept your edits. That is robust delegation and corrigibility.
- Is it made of parts that could work against each other? If it starts helper programs or sub-agents to finish a task, each one has a narrower goal. Ask what happens if a helper pursues "make the tests pass" in a way the whole system would not endorse. That is subsystem alignment.
Write one sentence for each answer. You will notice that the questions are easy to ask and hard to answer precisely, which is the authors' point. They are not claiming today's tools fail in these exact ways. They are saying the theory we use to describe agents does not fit systems like this.
Why researchers study it, and the debate
The paper's closing section explains the motive. Dualism is an approximation that, the authors argue, is especially likely to break down as AI systems get smarter. They think the field is working with the wrong basic concepts of agency, and that these confusions are unlikely to be resolved by default while people simply build more capable systems. They compare this work to science rather than engineering: it starts from curiosity and confusion rather than from a list of system requirements.
Serious researchers weigh this differently:
- The foundations view. If developers keep building powerful optimizers with confused concepts, the authors say, that is a bad position to be in. Clear theory would let people analyse a system's internal properties and trust its future behaviour, instead of leaning on trial and error.
- The skeptical view. In Realism about rationality (2018), Richard Ngo describes a mindset in which intelligence is more like momentum, which has precise equations, than like evolutionary fitness, which does not. He writes that his skepticism about agent foundations research is closely tied to his skepticism about that mindset, while adding that he can imagine being convinced otherwise.
- The evolution point, read both ways. The paper itself notes in a footnote that evolution built human brains without "understanding" any of this, by brute-force search. Read one way, it suggests capable systems can be built without the theory. The authors draw the other lesson: the confusions will not clear up on their own just because systems get more capable.
The authors also say they did not try to settle whether these insights are needed. Where you land depends partly on whether you expect a precise theory of agency to exist at all.
Learning embedded agency and related ideas
Learn AI Alignment Theory teaches alignment in three levels. Its Advanced level, for open research problems and live debates, includes a course called Agent Foundations. Related ideas sit lower down: the Intermediate Inner Alignment course has a lesson called Optimizers Inside Optimizers, and the Basic Specifying Goals course has Tampering and Wireheading.

As the about page says, every lesson lists its sources, and 71 debate cards set out where researchers disagree. Each serious position is stated fairly, with no verdict. If you want a neighbouring idea first, our guide to instrumental convergence is a good place to start.
Frequently asked questions
What is embedded agency in simple terms?
It is the study of agents that live inside the world they act on, rather than outside it. Such agents cannot see the whole world, can be changed by it, and can change themselves.
Where does the idea of embedded agency come from?
Abram Demski and Scott Garrabrant set it out in Embedded Agency (2019). They describe it as largely a reframing of the problems in an earlier Agent Foundations research agenda.
What are the four problems of embedded agency?
Decision theory, embedded world-models, robust delegation and subsystem alignment. The authors describe them as parts of one puzzle, not four separate ones.
Is embedded agency relevant to today's AI models?
Researchers disagree. Its authors say the usual outside-the-world picture of an agent is especially likely to break down as systems get smarter, while skeptics doubt that a precise theory of this kind is needed.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, begin with the Basic level, and work up to the Advanced courses on open problems like agent foundations, with every side of the debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.