Linear probes explained: reading what an AI model knows

Linear probes are one of the simplest tools for looking inside an AI model. A probe is a small classifier that tries to read one property, such as "is this statement true?", straight from the numbers inside a network. If a straight-line rule can read it out, the information is there in a form that is easy to get at.

This guide explains how linear probes work, walks through building a truth probe step by step, and covers what probes have found and why researchers argue about what they prove.

What linear probes are

When a neural network processes an input, each layer produces a long list of numbers called activations. You cannot read these lists by eye. A probe is a separate model you train to turn them into a label you care about.

The method was set out in Understanding intermediate layers using linear classifier probes by Guillaume Alain and Yoshua Bengio (2016). They trained linear classifiers on each layer of two image models, "trained entirely independently of the model itself", to see how useful each layer's features were. One finding: the deeper the layer, the easier it was to separate the classes with a straight line.

"Linear" is the important word. A linear probe can only draw a flat boundary: it multiplies each activation by a weight, adds them up, and compares the total with a threshold. That keeps it too simple to do much thinking of its own, which matters later.

Probes sit alongside the other tools in our guide to mechanistic interpretability for beginners.

How a linear probe works, step by step

  1. Pick a property. Something you can label, such as true or false, a part of speech, or whether a prompt will lead to harmful behavior.
  2. Collect labeled inputs. Hundreds or thousands of examples with the right answer attached.
  3. Record activations. Run each input through the model, which stays frozen, and save the activations at one layer.
  4. Train the probe. Fit a simple classifier, such as logistic regression, to predict the label from the activations.
  5. Test on new examples. Score the probe on inputs it never saw. High accuracy means the property can be read from that layer.
Diagram of a linear probe: labeled inputs go through a frozen model, activations at one layer feed a linear classifier, followed by three checks: control task, transfer and causal test

A worked example: a truth probe on paper

Here is how a truth probe study is set up. You can follow it without running anything.

  1. Write simple statements. For example, "Paris is in France" (true) and "Paris is in Japan" (false). Make a few hundred of each.
  2. Read one layer. Feed each statement to the model and keep the activations at the last word, from a middle layer.
  3. Take the difference of means. Average the activations for all true statements, then for all false ones. The line between the two averages is your probe's direction. A new statement is scored by how far it sits along that line.
  4. Check transfer. Test the probe on a different set, such as "seven is larger than three" style comparisons. If it still works, it is less likely to be tracking something about cities.
  5. Check cause. Push a false statement's activations along the direction and see whether the model now treats it as true.

This is close to what Samuel Marks and Max Tegmark did in The Geometry of Truth (2023). They found clear linear structure, probes that transferred across datasets, and interventions that made a model "treat false statements as true and vice versa". They also found that simple difference-of-means probes generalized as well as other methods while picking directions "more causally implicated in model outputs".

What linear probes have found

Beyond truth, probes have been pointed at questions that matter for safety.

  • Knowledge a model does not say. In Eliciting Latent Knowledge from Quirky Language Models (2023), Alex Mallen and colleagues fine-tuned models to answer wrongly whenever the name "Bob" was in the prompt. Linear probes, "especially in middle layers", usually reported what the model knew regardless of what it said.
  • Sleeper agents about to misbehave. Anthropic's note Simple probes can catch sleeper agents (April 2024) used models trained to act normally until a trigger, then do something harmful, such as writing insecure code. Probes built from generic question pairs like "Are you doing something dangerous? yes" and "... no" reached AUROC above 99% at predicting when the model would defect. AUROC is a score where 50% is chance and 100% is a perfect split. The probes knew nothing about the trigger or the harmful behavior.

Those sleeper agents are deliberately built stand-ins for the worry described in our post on deceptive alignment and scheming.

The catch: is the probe reading the model or doing the task?

A probe that scores well tells you something can be read from the activations. It does not tell you the model uses it, or that the probe found it rather than worked it out.

John Hewitt and Percy Liang put this sharply in Designing and Interpreting Probes with Control Tasks (2019). They asked whether high accuracy means that "the representations encode linguistic structure or just that the probe has learned the linguistic task". Their fix is a control task: give each word a random label and see how well the probe learns that. A good probe should be selective, scoring high on the real task and low on the random one. They found that popular probes on one well-known model were not selective, and that dropout did not make their larger probes more selective, though other forms of regularization did.

Yonatan Belinkov's review, Probing Classifiers: Promises, Shortcomings, and Advances (2021), collects these and other limitations of the method. Even Marks and Tegmark open by noting that truth probing "is controversial, with some authors pointing out failures of these probes to generalize in basic ways".

This is part of why researchers also use sparse autoencoders, which look for features without picking the label first, and why superposition makes any single direction harder to interpret.

Where researchers disagree

  • Probes are a practical safety tool now. Anthropic's note says classifiers like its defection probes "could form a useful part of AI control setups". They are cheap: they read the prompt's last token and need no sampling from the model.
  • The good results may not carry over. The same note says the strong signal may be partly due to how the sleeper agents were made, and that whether it holds for naturally arising deception "remains an important open question".
  • Reading is not the same as using. Hewitt and Liang argue that accuracy alone can mislead and that probes need controls. Marks and Tegmark add causal tests, to show a direction actually changes what the model outputs.

Frequently asked questions

What is a linear probe in machine learning?

A linear probe is a simple classifier trained on a model's internal activations to predict one property. It is trained separately, so the model itself is not changed.

Why use a linear probe instead of a bigger classifier?

A bigger classifier can learn the task by itself, so high accuracy would say little about the model. A linear probe is too simple to do much on its own, so its success says more about what the activations already contain.

Can linear probes detect when an AI is lying?

In test settings, probes have read a model's knowledge when its outputs were wrong on purpose, and have flagged sleeper agents before they misbehaved. Whether this works on deception that arises naturally is still an open question.

What is a control task for a probe?

It is a task with random labels that only the probe can learn. If a probe scores well on it too, its score on the real task may reflect the probe rather than the model.

Get started

Probes are one of several ways researchers try to see what a model represents. In Learn AI Alignment Theory, Interpretability is one of the Advanced courses, and the app has a glossary of 349 key terms. Every lesson lists its sources, and the debate cards set out each side with no verdict. You sign in with Google or an emailed code; the about page shows what is inside.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.