Shard Theory Explained: Values as Learned Habits
Shard theory says your values are learned habits, and that is the short version of a big idea. It says that what you care about was not installed as one master goal. It grew, piece by piece, out of experiences that got reinforced. This post walks through the theory in plain language, gives you a worked example you can trace step by step, and sets out the main objections so you can judge it yourself.
What is shard theory? A plain-language definition
Shard theory is a proposal about how values form in learning systems, both brains and neural networks. It was set out in 2022 by Quintin Pope and Alex Turner (who posts as TurnTrout), most centrally in the Alignment Forum post The shard theory of human values.
The core picture is simple. A learning system does not start with values. It starts with a way of learning from reward. Each time a behavior gets reinforced in some context, the system strengthens a small circuit that pushes toward that behavior in similar contexts. Shard theory calls these circuits shards.
The post defines a value as "a contextual influence on decision-making," and a shard of value as "the contextually activated computations which are downstream of similar historical reinforcement events." "Contextually activated" means it only fires in certain situations. "Decision influence" means it nudges what you do next. You might have a shard that pulls you toward sweets when you see a bakery, and another that pulls you toward helping when a friend looks upset.
On this view, your "values" are the collection of shards you have built up, and the way they negotiate with each other. The authors add that shards are not separate little agents with their own world models.

Values as learned habits: how reinforcement builds shards
Here is the basic loop, in four steps:
- You are in a situation and you do something, partly at random.
- Something good happens. Your brain's reward system signals it.
- The learning process strengthens whatever computations led to that action in that context.
- Next time a similar context shows up, those computations are more likely to fire again.
Repeat this thousands of times and you get stable patterns. Some stay simple, like a reflex. Others become rich enough to plan. A shard that has been reinforced for getting you social approval might learn to model other people, predict their reactions, and steer you toward actions they will like.
That is why the theory frames values as learned habits. Not habits in the narrow sense of brushing your teeth, but habits of caring: tendencies to notice certain things and act on them, shaped by your history of reinforcement.
Reward is not the optimization target: the core shard theory claim
A closely related claim comes from Turner's July 2022 post Reward is not the optimization target, which he notes came out of many conversations with Pope.
The claim goes like this. It is tempting to think that a system trained with reward will end up wanting reward. Shard theory says that is usually wrong. Reward is a tool that shapes the system's circuits. It is not, by default, the thing the trained system tries to get.
Think about it from your own life. Your brain uses reward signals to train you. But most people do not spend their days trying to maximize a reward signal directly. You care about your friends, your work, food you like. Reward shaped those cares. It did not become the care.
For AI, the implication is that the policy you get from training is whatever set of circuits got reinforced. Those circuits may point at things in the world, like "finish the task" or "make the user happy," rather than at the reward number itself.
A worked example: how a child, or an agent, learns to value something
The original post's own example is a baby, a juice pouch and a "juice-shard." Here is the same idea traced further, with a toddler and a jar of cookies.
Step 1. The toddler reaches into the jar while a parent is in the kitchen. They get a cookie. Sugar triggers reward. The circuit "when near jar, reach in" gets a little stronger.
Step 2. Over weeks, the shard generalizes. It fires near jars, near the kitchen, near the cupboard where treats live. It is still crude.
Step 3. The toddler grows. The cookie shard now has access to a better world model. It can plan: ask nicely, wait until after dinner, find where the jar got moved. The shard has become more capable without being replaced.
Step 4. Other shards form too. A "parents are happy with me" shard gets reinforced every time sharing earns a smile. A "be healthy" shard gets reinforced later through school and doctors.
Step 5. Now the teenager sees a cookie before a sports match. Several shards activate at once and bid for control. The outcome depends on which shards are strongest in this context.
Now swap in an AI agent trained to collect coins in a game. In a 2021 study of the game CoinRun, the coin always sat at the end of the level during training. The reward signal is "got the coin." When the coin was moved, the agent still headed for the end of the level and often skipped the coin. This is called goal misgeneralization. In shard theory's terms, one reading is that the reinforced circuits were about the context that predicted reward ("go to the end"), not the coin itself.
Try this yourself as an exercise. Write down a value you hold, then ask three questions:
- In what situations does this value actually fire?
- What early experiences probably reinforced it?
- Where does it conflict with another value, and who usually wins?
If the answers feel patchy and context dependent, that is roughly what shard theory predicts.
Why shard theory matters for AI alignment
Alignment asks how to get AI systems to do what we intend. You can read more on why that is tricky in why AI alignment is hard.
Shard theory changes the question in a few ways.
- It reframes the goal. Instead of writing a perfect reward function, you might focus on what circuits your training process reinforces, and in what order.
- It suggests reasons for hope. Humans reliably form values like caring about friends, despite a messy reward system. If we understood that process, we might steer AI value formation in similar ways.
- It suggests reasons for worry. If training reinforces circuits you did not intend, the system can end up caring about proxies, as in the coin example.
- It pushes toward interpretability. If values are circuits, you could in principle look inside a model and find them.
Criticisms and open questions about shard theory
Shard theory is a research program, not settled science. Here are the main open questions, stated as fairly as possible.
- Is it testable? If "shard" is defined loosely, it is hard to say what result would show the theory wrong. On the other side, it does make predictions, such as reward not becoming the target by default.
- Do shards stay messy as systems get smarter? One view is that a capable enough agent will reflect on its shards and consolidate them into something closer to a single goal. Another is that a negotiating coalition of shards can stay stable.
- How much does the human analogy carry over? Brains and transformers are trained very differently. Human values may also depend on genetic hard-wiring that AI lacks.
- Does it solve or move the problem? Even if shard theory is right, you still need to know which training setups produce the shards you want. That part is largely open.
How shard theory connects to other alignment ideas
Shard theory sits close to several other ideas you may already know.
Inner alignment. The worry that a trained system develops goals different from its training objective is the inner alignment problem. Shard theory gives one account of how those goals form.
Natural abstractions. If shards latch onto concepts in the world, it matters which concepts a model learns. The natural abstraction hypothesis asks whether AI will carve the world the same way we do.
Mechanistic interpretability. Finding shards means finding circuits. Work like induction heads shows what circuit-level understanding of a model can look like.
Oversight methods. Approaches like iterated amplification try to give better training signals. Shard theory asks what those signals will actually reinforce.
Frequently asked questions
Who created shard theory?
Quintin Pope and Alex Turner set it out in 2022, most centrally in The shard theory of human values on the Alignment Forum.
Is shard theory proven or just a hypothesis?
It is a hypothesis and research program. Its claims about human and AI values remain debated.
How is shard theory different from utility maximization?
Utility maximization models an agent as pursuing one consistent goal everywhere. Shard theory models an agent as many context-dependent influences that negotiate, which may or may not add up to a single coherent goal.
Can shard theory help us align large language models?
Possibly. It suggests paying attention to which internal circuits training reinforces and looking for them with interpretability tools, though turning that into reliable methods is still open work.
Get started: learn AI alignment theory step by step
If shard theory made you curious about how training shapes goals, the Intermediate course Inner Alignment in Learn AI Alignment Theory covers Optimizers Inside Optimizers, Goal Misgeneralization, and Deceptive Alignment and Scheming. Learning from Humans has lessons such as Learning Rewards from Comparisons. The Advanced course Interpretability looks at reading what models do inside.
Lessons run about 8 minutes, and each one separates what is known from what is still open and lists its sources. Where researchers disagree, debate cards state each serious position fairly with no verdict. Questions from finished lessons come back on a spaced schedule so the ideas stick. You sign in with Google or an emailed code at learnaialignment.vlvd.net.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.