Linear probes are one of the simplest tools for looking inside an AI model. A probe is a small classifier that tries to read one property, such as "is this statement true?", straight from the numbers inside a network. If…
Interpretability
4 posts
Linear probes explained: reading what an AI model knows
Superposition in neural networks: why neurons mix ideas
Superposition in neural networks is the idea that a model can store more concepts than it has neurons, by overlapping them. It is one reason a single neuron in a large model often responds to several unrelated things,…
Sparse autoencoders explained: finding features in AI models
Sparse autoencoders are a tool for reading what is going on inside an AI model. A model's neurons each mix many ideas together, so looking at one neuron tells you little. Sparse autoencoders try to pull those mixed…
Mechanistic interpretability for beginners
A trained neural network does useful things, but nobody wrote down how. Its knowledge sits in millions or billions of numbers. Mechanistic interpretability is the research effort to read those numbers: to explain what a…