Superposition in neural networks: why neurons mix ideas
Superposition in neural networks is the idea that a model can store more concepts than it has neurons, by overlapping them. It is one reason a single neuron in a large model often responds to several unrelated things, and one reason reading a model's insides is so hard.
Our guide to sparse autoencoders covers a tool built to undo superposition. This guide is about the thing itself: why a network would do it, what small experiments showed, a pencil example you can check, and what is still unknown.
The puzzle: neurons that mean several things
If every neuron stood for one idea, you could read a model like a list. Researchers have found that many neurons do not work that way. One neuron can fire for several unrelated concepts. This is called polysemanticity: one neuron, many meanings.
The 2022 paper Toy Models of Superposition, by Elhage and colleagues at Anthropic and Harvard, set out to explain it. Its proposal: the model is storing extra features, more than it has room for one per neuron, and the overlap is what makes neurons look mixed.
A feature here means a property of the input the model tracks, such as "this text is in French". The paper treats features as directions in the space of a layer's activations, the numbers the layer holds for one input.
How superposition in neural networks works
With two neurons you have two dimensions, so only two directions can sit at right angles, where they never interfere. To store a third feature you have to point it somewhere that overlaps with the others. Then, when one feature is on, a little of it leaks into how the others read. The paper calls that leak interference.
Interference is only a problem when features are on at the same time. If features are sparse, meaning each one is rarely active and they seldom fire together, the model can afford overlap. A nonlinearity, such as the ReLU that sets negative values to zero, can filter out the small leaks.
So why would a feature ever line up with a single neuron? The paper names two forces pulling in opposite directions. Some layers have what it calls a privileged basis: something about how they are built, such as applying an activation function, can encourage features to align with individual neurons. Superposition pulls the other way, spreading features across neurons so more of them fit. Which force wins helps explain why some neurons are clean, responding to one feature, while others are mixed. In the toy models, both kinds formed.
What the toy models showed
The authors trained very small networks on made-up data where they knew the true features. In one setup, five features of different importance had to fit into two dimensions.

- Dense features. When features were often on together, the model kept the two most important at right angles and did not store the other three.
- Sparse features. When features were rare, the model stored more of them, overlapping, and accepted the interference.
- A phase change. Whether a feature is stored in superposition switched suddenly, not gradually, as sparsity and importance changed. The paper uses the term for a discontinuous change.
- Geometry. Features arranged themselves into regular shapes such as triangles and pentagons.
- Computation. In limited cases the models could also compute while in superposition, not only store.
The paper suggests real models might be noisily simulating much larger, sparser networks. It also reports preliminary evidence of a link to adversarial examples, the small input changes that fool a model.
A worked example: interference with a pencil
Here is a made-up model with two neurons and three features, A, B and C, spread evenly at 120 degrees. Their directions are A = (1, 0), B = (-0.5, 0.87) and C = (-0.5, -0.87). To read a feature, you multiply the activation by its direction and add, then apply ReLU.
- Only A is on. The activation is (1, 0). Read A: 1. Read B: -0.5. Read C: -0.5.
- Filter. ReLU turns the negatives to 0. You get A = 1, B = 0, C = 0. A perfect read, even with three features in two neurons.
- Now A and B are both on. The activation is their sum, (0.5, 0.87). Read A: 0.5. Read B: about 0.5. Read C: about -1.
- Filter again. You get A = 0.5, B = 0.5, C = 0. C is correctly off, but A and B each read at half strength.
Steps 1 and 2 show why sparse features suit superposition: one at a time, the overlap costs nothing after filtering. Steps 3 and 4 show the price when features fire together. It also shows why looking at one neuron misleads you: the first neuron's value here mixes A, B and C.
Where researchers disagree
- How far the toy results reach. The authors say their models are simple ReLU networks, and that it is very unclear what carries over to real networks. How much of what large models do it explains is a question for later work.
- Are concepts really directions? The paper builds on the linear representation hypothesis, the idea that concepts are directions in activation space. Park, Choe and Veitch (2023) report evidence of linear representations of concepts in LLaMA-2.
- Not every feature is a single direction. Engels and colleagues (2024) found features in GPT-2 and Mistral 7B that are inherently multi-dimensional, such as circles for the days of the week and the months, used to do arithmetic on them. They argue such features are needed to explain some behavior.
- What it means for safety. The toy models paper says that if superposition happens in real networks, it deeply shapes which approaches to interpretability make sense. That is part of the case for tools like sparse autoencoders, and part of why the field in our guide to mechanistic interpretability for beginners is hard.
Learn AI Alignment Theory sets out each of these positions with no verdict.
Learning interpretability step by step
Learn AI Alignment Theory has an Interpretability course in its Advanced level, the level for open research problems and live debates.

The Basic level needs no background, so you can start there and work up. The about page shows what you do in each lesson, and every lesson lists its sources.
Frequently asked questions
What is superposition in a neural network?
It is a model storing more features than it has neurons by giving them overlapping directions. It works best when those features are rarely active at the same time.
Is superposition the same as polysemanticity?
No. Polysemanticity is what you observe, one neuron responding to several concepts. Superposition is a proposed cause of it.
Does superposition happen in large language models?
It was shown clearly in small toy models. Whether and how much it explains large models is still being studied.
How do researchers undo superposition?
One common tool is the sparse autoencoder, which tries to recover the overlapping features as separate directions. Our guide to sparse autoencoders explains how it works and its limits.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, work up from the Basic level to the Interpretability course, with every side of the debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.