Mechanistic interpretability for beginners

A trained neural network does useful things, but nobody wrote down how. Its knowledge sits in millions or billions of numbers. Mechanistic interpretability is the research effort to read those numbers: to explain what a model does in terms of its internal parts, the way you might explain a machine by tracing its gears.

This guide is for beginners. It covers the core ideas without math, shows one hands-on way to think about the work, and sets out why researchers disagree about how much it can do for AI safety.

What mechanistic interpretability means

Most ways of studying AI look at inputs and outputs: what goes in, what comes out. Mechanistic interpretability looks inside. In the words of one well-known paper, it seeks to explain behaviors of machine learning models in terms of their internal components.

Why this matters for safety: if you could read what a model represents and how it reaches an answer, you might spot a problem that never shows up in its behavior during testing. Our post on deceptive alignment and scheming explains the kind of hidden problem people have in mind.

The first obstacle: neurons that mean many things

The natural place to start is the single neuron. Sometimes a neuron does seem to stand for one thing. Often it does not. Anthropic's Towards Monosemanticity (2023) gives a famous example from an image model, Inception v1: one neuron responds both to cat faces and to the fronts of cars.

Researchers call this polysemanticity, a neuron that responds to several unrelated things. It is a problem because you cannot explain a model one neuron at a time if each neuron is a mix.

Superposition: more ideas than neurons

Toy Models of Superposition (Elhage and colleagues, 2022) built small models where the cause can be seen fully. The models store more features, meaning more distinct ideas, than they have neurons. They do this by spreading each feature across many neurons as a direction, rather than giving each one its own neuron. The paper calls this superposition, and it shows that polysemantic neurons follow from it.

Diagram from a polysemantic neuron, to superposition, to a sparse autoencoder, to readable features such as Arabic script, DNA, base64 and Hebrew

If that is what is going on, the right unit to study is not the neuron but the direction.

Sparse autoencoders: pulling the ideas apart

A sparse autoencoder is a second network trained to rebuild a model's internal activity from a larger set of directions, with only a few active at once. The hope is that each of its directions picks out one feature.

Cunningham and colleagues (2023) found that sparse autoencoders on a language model learn features that are more interpretable and single-meaning than other methods find. Towards Monosemanticity applied the idea to a one-layer transformer with a 512-neuron layer, training on 8 billion data points and pulling out anywhere from 512 to 131,072 features. Among the features studied were ones for Arabic script, DNA, base64 and Hebrew.

Circuits: how parts work together

Features are the nouns; circuits are the sentences. A circuit is a group of parts that together carry out one behavior.

The best-known example is Interpretability in the Wild (Wang and colleagues, 2022). The authors explained how the small language model GPT-2 small does a task called indirect object identification, working out which name a sentence should end with. Their explanation involves 26 attention heads grouped into 7 classes, found by switching parts on and off and watching what changed. They tested it for faithfulness, completeness and minimality, and they report that those same tests point to remaining gaps in the explanation.

A worked example: checking one feature yourself

Here is a simplified version of how you would judge whether a feature from a sparse autoencoder is real. The feature below is made up for illustration.

  1. Find the texts that light it up most. Suppose the top 20 are all strings of letters, digits, plus signs and slashes ending in equals signs.
  2. Name a guess. "This feature responds to base64 text."
  3. Test the guess on new inputs. Feed in fresh base64 and fresh ordinary English. It should fire on the first and stay quiet on the second.
  4. Check the weaker activations too. If it also fires weakly on random code or hex strings, your name is too narrow or the feature is still a mix.
  5. Change it and watch. Turn the feature up or down inside the model and see whether its outputs shift the way your name predicts.

Steps 3 to 5 are what separate a story from an explanation. A label that only fits the top examples is a guess.

Where researchers disagree

  • The case for. The IOI paper presents its result as evidence that a mechanistic understanding of large models is feasible. Cunningham and colleagues hope sparse autoencoders can be a foundation for future work and for greater transparency and steerability.
  • The case against. In Against Almost Every Theory of Impact of Interpretability (2023), Charbel-Raphaël argues that the overall case for how interpretability will make AI safer is weak, that studying today's models may not predict future ones, and that using interpretability to audit a model for deception is out of reach. He still calls interpretability research commendable.
  • The open question. Both sides agree today's results come from small models or narrow tasks. Whether the methods scale to the largest systems is unknown.

Learn AI Alignment Theory sets out these positions side by side, with no verdict.

Learn interpretability step by step

Learn AI Alignment Theory has an Interpretability course in its Advanced level, one of seven Advanced courses on open research problems and live debates.

The Three levels panel on the Learn AI Alignment Theory about page: Basic, Intermediate and Advanced, with how many courses each has

Every lesson lists its sources, and 71 debate cards set out where researchers disagree. For a route from the basics to this level, see our AI safety self-study path, or the about page for all 22 courses.

Frequently asked questions

What is mechanistic interpretability in simple terms?

Reverse-engineering a trained neural network: explaining what it does in terms of its internal parts, such as features and circuits.

What is superposition?

When a network stores more features than it has neurons by spreading each feature across many neurons, which makes single neurons hard to read.

What does a sparse autoencoder do?

It rebuilds a model's internal activity from a larger set of directions, few active at once, so that each direction is more likely to stand for one idea.

Do I need to code to learn it?

Not to understand the ideas. Research work uses code, but the concepts above can be learned first on their own.

Get started

Start Learn AI Alignment Theory: sign in with Google or an emailed code, build up through the Basic and Intermediate levels, then take the Interpretability course, with every side of the debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.