Logit Lens Explained: Reading a Model Layer by Layer
Here is the logit lens explained in short: a language model makes its final guess at the very last layer. The logit lens lets you ask the model what it would guess at every earlier layer too. You take the hidden state from the middle of the network and push it through the model's own output layer. The result is a rough picture of how the prediction forms, one layer at a time. This post shows how it works, walks through an example on GPT-2, and explains where the method fails.
Logit lens explained: a plain definition
A transformer turns your text into vectors, changes those vectors over many layers, and then converts the final vector into a score for every token in its vocabulary. Those scores are called logits. The highest one is the model's top guess for the next token.
The logit lens is a simple trick. It applies that same final conversion to vectors from earlier layers. You do not train anything new. You reuse the model's own "unembedding" step and read off what each layer seems to be predicting.
It is one of the cheapest tools in interpretability, because it needs no training and only a few lines of code. Belrose and colleagues, who later improved on it, describe it as a technique that "yielded useful insights but is often brittle".
How a transformer builds a prediction across layers: the residual stream
To see why the trick works at all, you need one concept: the residual stream.
Think of the residual stream as a shared notepad that runs through the whole model. Each token gets its own vector on this notepad. Every layer reads the current vector, computes something, and adds its result back in. Attention layers move information between token positions. MLP layers transform information at a single position. Neither one replaces the vector. They only add to it.
That additive design matters. Because each layer writes into the same space, the vector at layer 6 lives in roughly the same "language" as the vector at layer 12. So it is reasonable to hope that the output layer, which was built to read the final vector, can also read a middle one. The logit lens tests that hope.
How the logit lens works step by step: unembedding intermediate layers
Here is the full procedure for one prompt.
- Run the model and save hidden states. Feed in your prompt and record the residual stream vector after each layer, at the position where the next token will be predicted (usually the last token).
- Apply the final layer norm. Most transformers normalize the last vector before the output layer. Apply that same normalization to each saved vector.
- Multiply by the unembedding matrix. This turns each vector into one logit per vocabulary token.
- Convert to probabilities. Apply softmax so you can compare layers on the same scale.
- Read the top tokens per layer. List the top few tokens, or track the probability of one target token, from the first layer to the last.
The output is often drawn as a grid. Rows are layers, columns are token positions, and each cell shows the top predicted token. You read it from bottom to top and watch the guesses change.
Worked example: watching GPT-2 settle on the next token layer by layer
Take the prompt The Eiffel Tower is in the city of. The answer you expect is Paris.
In any library that lets you cache a model's activations, the steps are the same: run the prompt, collect the residual stream after each layer at the last position, apply the final layer norm, multiply by the unembedding matrix, and apply softmax. Then list the top tokens for each layer.
What should you look for? The numbers below are not measurements. They describe a shape to look for, so you know what to check in your own run.
- Early layers: the top guesses are often just echoes of the current input token, or common filler tokens like
theor,. The probability ofParisis near zero. - Middle layers: related tokens start to show up. You might see other city names, or words tied to France. This hints that the model has pulled up "location" information before it has picked a specific answer.
- Later layers:
Parisclimbs to the top and its probability rises sharply. The final layers mostly sharpen a choice that was already made.
Run this yourself and note the layer where Paris first becomes the top token. Then change the prompt to something harder, like The capital of the country north of Spain is. Does the answer show up later? Treat what you see as a clue, not proof.

Where the logit lens breaks down and why the tuned lens fixes some of it
The logit lens rests on one assumption: middle layers use the same coordinate system as the last layer. That assumption is often wrong.
- Early layers can be unreadable. The information may be there, but stored in a form the output layer cannot read.
- Representations drift. Each layer can rotate or rescale features a little. Small drifts add up, so a lens built for the last layer fits earlier layers poorly.
- It shows a guess, not a reason. Seeing
Parisat layer 8 tells you the vector points toward Paris. It does not tell you which attention heads or MLPs put it there.
The tuned lens, from Nora Belrose and colleagues in 2023, tackles this. It trains a small affine probe for each block of a frozen model, so every hidden state can be decoded into a prediction over the vocabulary. Tested on language models up to 20 billion parameters, they report it is more predictive, reliable and unbiased than the logit lens, and causal experiments suggest it uses features similar to the model's own.
The tradeoff is that you now have a trained component. A learned translator might add structure that was not really in the model. So the tuned lens is more reliable for reading predictions, but you should still check findings with other methods, such as interventions that change an activation and measure the effect.
Why the logit lens matters for AI alignment, evals, and detecting hidden behavior
For alignment, a useful question is not only "what did the model output?" but "what was it computing along the way?" The logit lens gives a first look, alongside tools such as induction heads research in mechanistic interpretability.
Here are a few ways researchers think about using lenses like this:
- Checking whether the model "knew" something it did not say. If the correct answer is the top guess at a middle layer and then gets pushed down by the final layers, that is worth looking into. It connects to the open problem of eliciting latent knowledge.
- Supporting evaluations. Behavioral tests only see outputs. Internal reads can add a second line of evidence. For more on what tests measure, see AI model evaluations explained.
- Looking for hidden behavior. Belrose and colleagues found that the path of a model's layer-by-layer predictions can detect malicious inputs with high accuracy. Work on models trained with backdoors asks a similar question. The background is in sleeper agents in AI explained.
- Judging reasoning, not just answers. Seeing when an answer forms is loosely related to the idea of supervising process rather than results, covered in process supervision vs outcome supervision.
Researchers disagree about how far tools like this can go. Some argue that reading internal states is key to catching deception. Others argue that a capable model's internals may be too complex for simple projections to be trusted. Both views deserve a fair hearing. The logit lens alone settles neither one.
Frequently asked questions
Is the logit lens reliable?
Only partly. Its own successors call it often brittle, which is why the tuned lens adds a small trained map per layer before decoding.
What is the difference between the logit lens and the tuned lens?
The logit lens decodes each layer directly with the model's final layer norm and unembedding. The tuned lens first passes each layer through a small learned map trained to match the final output, which usually gives cleaner early-layer readings.
Does the logit lens work on any model?
It runs on any transformer with a residual stream and an output layer, but how readable the middle layers are varies. The tuned lens was tested on models up to 20 billion parameters.
How is it different from a linear probe?
A logit lens reuses the model's own output layer, so it can only read what maps onto next-token predictions. A linear probe is trained to read any property you label.
Get started
The logit lens makes more sense once you understand why reading a model's internals matters in the first place. Learn AI Alignment Theory builds that context in short lessons of about 8 minutes. The Advanced level includes an Interpretability course and a Scalable Oversight course with a lesson on latent knowledge. The Intermediate Inner Alignment course covers goal misgeneralization and deceptive alignment and scheming, the failures interpretability tools are often meant to catch.
Each lesson separates what is known from what is still open and lists its sources. Debate cards set out where researchers disagree, with every serious position stated fairly and no verdict. Questions from lessons you finish come back on a spaced schedule so the ideas stick. A glossary of 349 terms and a topics map that follows the aisafety.com self-study topics help you see where interpretability fits. You sign in with Google or an emailed code. Read more on the about page.
If you are new to the field, the AI alignment reading list for beginners is a good companion.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.