Sparse autoencoders explained: finding features in AI models
Sparse autoencoders are a tool for reading what is going on inside an AI model. A model's neurons each mix many ideas together, so looking at one neuron tells you little. Sparse autoencoders try to pull those mixed signals apart into separate features, each ideally standing for one concept a person can name.
Our guide to mechanistic interpretability for beginners introduces the field. This guide goes one level deeper on one of its main tools: how it works, what it found in large models, and why researchers disagree about how much it can do for safety.
The problem sparse autoencoders solve
You might hope each neuron in a model stands for one thing. Often it does not. Toy Models of Superposition (Elhage and colleagues, 2022) describes neurons that respond to many unrelated concepts, which it calls polysemanticity, and shows in small models how it can arise. The model stores more features than it has neurons by overlapping them, a pattern the paper calls superposition.
If concepts are packed in overlapping directions rather than in single neurons, you need a way to find those directions. That is the job of a sparse autoencoder.
How a sparse autoencoder works
A sparse autoencoder is a second network trained on the inside of the first. It takes the activations from one layer of the model, the list of numbers the model holds at that point, and rebuilds them from a much larger set of candidate features.
The key word is sparse. For any one input, only a handful of those features may be switched on. Cunningham and colleagues (2023) trained sparse autoencoders on a language model's activations and found the learned features were more interpretable, and more often had a single meaning, than directions found by other methods.

Training balances two pulls. The rebuilt activations should match the originals, and as few features as possible should be on. Push too hard on either and the other suffers.
The same paper tested whether the features do real work. On a known language task, where a model has to pick out the right name in a sentence, the features let the authors pinpoint which parts of the model were causally responsible for the answer, more precisely than earlier ways of splitting up the activations. Finding a feature is one step; showing that it matters to the model's behavior is the stronger claim.
A worked example: why sparsity picks one explanation
You can see the main idea with two neurons and a pencil. The numbers and labels below are made up for illustration.
- Set up the features. A tiny model has two neurons. A sparse autoencoder has learned three feature directions: "French text" is (1, 0), "code" is (0, 1) and "bridges" is (0.6, 0.8).
- Read one activation. On some input, the two neurons read (0.6, 0.8).
- Find explanations that rebuild it. Explanation A: "bridges" on at strength 1. That gives (0.6, 0.8) exactly. Explanation B: "French text" at 0.6 plus "code" at 0.8. That also gives (0.6, 0.8) exactly.
- Apply the sparsity rule. Both rebuild the activation perfectly, but A uses one feature and B uses two. A sparse autoencoder is trained to prefer A.
- Check the label. Collect the inputs where "bridges" is most active. If they are about bridges, the label holds up. If they are a mix, it was a guess.
Step 3 is superposition in miniature: two neurons, three concepts, and more than one way to read the same numbers. Step 4 is the bet sparse autoencoders make, that the simplest reading is the real one.
What scaling them up found
In 2024, Anthropic reported mapping millions of features in the middle layer of Claude 3.0 Sonnet. The page calls the method dictionary learning; the full paper describes it as scaling up sparse autoencoders. One feature responded to the Golden Gate Bridge across many languages and in images. Turning it up artificially made the model describe itself as the bridge. Turning up a feature linked to scam emails overcame the model's harmlessness training and it drafted one.
The team also reported features tied to safety concerns, including code backdoors, bias, power-seeking, manipulation, secrecy and sycophantic praise. Turning features up changed the model's behavior, which the team takes as a sign that features shape what the model does rather than only matching the input text.
Around the same time, Gao and colleagues (2024) studied how sparse autoencoders scale. They note that balancing rebuilding against sparsity is hard to tune and that many features can end up never used, called dead latents. They proposed a version that fixes how many features may be on at once, found clean scaling laws, and trained a sparse autoencoder with 16 million latents on GPT-4 activations.
Where researchers disagree
- The promising view. Anthropic suggests these features might one day be used to monitor models for dangerous behaviors, such as deceiving the user, and to steer them. That links to worries in our guide to deceptive alignment and scheming.
- The limits the builders name. The same Anthropic page says the features found are a small subset of what the model learned, and finding a full set with current methods would cost more compute than training the model. Knowing the features also does not show how the model uses them, and their safety value still has to be shown.
- The skeptical view. Kantamneni and colleagues (2025) point out there is no ground truth for which concepts a model really uses. They tested sparse autoencoders on practical probing tasks and found they did not consistently beat simple baselines. They call for testing interpretability methods on real tasks against strong baselines.
Learn AI Alignment Theory sets out each of these positions with no verdict.
Learning sparse autoencoders step by step
Learn AI Alignment Theory has an Interpretability course in its Advanced level, the level for open research problems and live debates. Starting from the Basic level, which needs no background, gives you the groundwork first.

The glossary explains 349 key terms, and every lesson lists its sources. The about page shows what you do in each lesson, and our AI safety self-study path shows where interpretability fits.
Frequently asked questions
What are sparse autoencoders in simple terms?
They are extra networks that rebuild a model's internal activity from a large set of features, with only a few on at a time. The goal is features that each stand for one idea.
Why do they need to be sparse?
Many combinations of features can rebuild the same activity. Keeping few features on picks the simplest reading, which is more likely to match single concepts.
Can sparse autoencoders find every concept in a model?
Not today. Anthropic says the features it found are a small subset, and a full set would cost more compute than training the model.
Are sparse autoencoders proven useful for safety?
Not yet. Their builders say safety use still has to be shown, and a 2025 study found they did not consistently beat simple baselines on probing tasks.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, work up from the Basic level to the Interpretability course, with every side of the debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.