Induction Heads Explained: How Language Models Copy Text

If you have ever watched a language model repeat a strange name or a made-up code it saw earlier in a prompt, you have seen induction heads at work. This guide explains induction heads: how language models copy text from earlier in their context, step by step. You will learn the two-head circuit behind it, the sudden jump during training when it forms, and why interpretability researchers care so much about it.

What are induction heads? A plain-English definition

An induction head is a specific kind of attention head inside a transformer. Its job is simple to state. When the model sees a token it has seen before, the induction head looks back to that earlier spot, finds the token that came right after it, and pushes the model to predict that same token again.

Put another way: "Last time I saw this, that came next. So predict that again." The word "induction" comes from this guess that a pattern will repeat. It has nothing to do with mathematical proof by induction.

Researchers at Anthropic named induction heads in 2021 in A Mathematical Framework for Transformer Circuits (Elhage and colleagues). They studied them further in 2022 in In-context Learning and Induction Heads (Olsson and colleagues).

How language models copy: the [A][B] ... [A] -> [B] pattern

Take this prompt:

The inventor Zorvath Klimbe built a clock. Years later, people still talk about Zorvath

"Zorvath Klimbe" is a made-up name, so a model cannot be recalling a fact about it. If it predicts "Klimbe" next, it is copying from earlier in the prompt.

The pattern is written as [A][B] ... [A] -> [B]. Here [A] is "Zorvath" and [B] is "Klimbe". The first time the model sees [A] followed by [B], nothing special happens. When [A] shows up again, the induction head finds the earlier [A], moves one step forward to [B], and copies [B] into the prediction.

This works on pure nonsense too. The 2022 paper actually defines induction heads by how they behave on repeated copies of random token sequences, so that meaning cannot be doing the work. It is just the pattern.

Diagram of the induction head circuit: a previous-token head labels Klimbe as following Zorvath, then the induction head at the second Zorvath finds that label and predicts Klimbe

The two-head circuit: previous-token heads and induction heads working together

One attention head cannot do this alone. Here is the problem. When the model sits at the second "Zorvath", it wants to attend to "Klimbe". But "Klimbe" holds no clue that it came after "Zorvath". Each token, by default, mostly carries information about itself.

So the circuit uses two heads in different layers:

  1. The previous-token head (earlier layer). At each position, it attends to the token just before and copies some of that token's identity into the current position. After this step, the position holding "Klimbe" also carries a note that reads roughly "the token before me was Zorvath."
  2. The induction head (later layer). At the current "Zorvath", it searches for any position whose note says "the token before me was Zorvath." It finds "Klimbe", attends there, and copies "Klimbe" forward as the prediction.

The first head writes a label. The second head searches by that label. Neither one is enough without the other.

Inside the mechanism: QK matching, OV copying, and K-composition

Each attention head has two jobs, and the Elhage paper splits them into two circuits.

  • The QK circuit (query and key) decides where to look. The query comes from the current token. Each earlier position offers a key. A high match means high attention.
  • The OV circuit (output and value) decides what to move once the head has looked. For an induction head, the OV circuit copies: attending to "Klimbe" raises the odds of outputting "Klimbe".

The key trick is called K-composition. The induction head builds its keys from information the previous-token head already wrote into the residual stream. So the key at "Klimbe" encodes "preceded by Zorvath". The query at the current position encodes "I am Zorvath". They match, and the head attends to the right place.

This explains a hard limit. A one-layer, attention-only transformer has no earlier head to compose with. The 2021 paper reports that induction heads "only develop in models with at least two attention layers."

The phase change: how induction heads suddenly form during training

Induction heads do not appear slowly. In the 2022 paper, they show up in a sudden window early in training. Several things happen together during that window:

  • The training loss curve shows a visible bump, a small plateau and then a drop.
  • Heads with induction behavior appear in the model.
  • The model gets much better at using tokens far back in its context.

The authors measured that last change with an "in-context learning score": the loss on the 500th token of a context minus the loss on the 50th. In-context learning develops in a narrow window, roughly 2.5 to 5 billion tokens into training, growing from under 0.15 nats to about 0.4 nats, and then stays constant for the rest of training.

The paper reports this phase change for language models of every size it studied, as long as they have more than one layer. That is part of what made the result striking. A neat mechanism and a sharp change in behavior lined up in time.

Induction heads and in-context learning: what the evidence shows

The bold claim in the 2022 paper is that induction heads may account for most in-context learning in transformers. The authors were careful about how strong their evidence was, and you should be too.

Here is what is fairly well supported:

  • Small attention-only models: the authors had strong, near-causal evidence. They could remove the heads and watch in-context learning collapse.
  • Timing: in larger models, induction heads and the jump in in-context learning appear at the same point in training.
  • Architecture changes: a "smeared key" tweak makes the previous-token step easy. With it, a one-layer model gets the phase change, and a two-layer model gets it earlier, as the theory predicts.

Here is what is still open:

  • Large models with MLP layers: the evidence was mostly correlational. Timing that matches is not the same as proof of cause.
  • Abstract tasks: few-shot reasoning and translation may use richer circuits that only partly resemble simple copying.

Keeping known and open separate matters here. A tidy story about one circuit can tempt you to stretch it further than the data supports.

Why induction heads matter for AI alignment and interpretability

Induction heads are one of the clearest cases where researchers took a real behavior, found the parts that produce it, and explained how those parts fit together. That is the core promise of mechanistic interpretability.

For alignment, this matters for a few reasons:

  • Proof of concept. It shows that neural networks are not always a fog of numbers. Some behaviors have readable mechanisms.
  • Training surprises. The phase change shows that capabilities can appear suddenly. If safety-relevant behaviors also arrive in jumps, gradual monitoring could miss them.
  • Debate about scaling. One reading treats induction heads as early evidence that interpretability can scale to frontier models. Another notes that a two-layer copying circuit is far simpler than the deceptive or goal-directed behavior safety work worries about. Both readings deserve a hearing.

If you are new to the field, start with mechanistic interpretability for beginners. A harder case than induction heads is superposition, where one neuron mixes several ideas.

Interpretability also links to other safety ideas. Tools that let humans check what a model is doing support oversight methods like iterated amplification.

Frequently asked questions

Do induction heads exist in one-layer transformers?

No, not in a standard one-layer, attention-only transformer. The circuit needs a previous-token head in an earlier layer for the induction head to compose with, so it takes at least two layers.

Are induction heads the same as in-context learning?

No. The 2022 paper argues induction heads may drive most in-context learning. The evidence is causal in small attention-only models and correlational in larger ones.

Who discovered induction heads?

Researchers at Anthropic named them in "A Mathematical Framework for Transformer Circuits" (Elhage and colleagues, 2021). They studied the tie to in-context learning in "In-context Learning and Induction Heads" (Olsson and colleagues, 2022).

Can induction heads do fuzzy or abstract copying?

The 2022 paper reports heads in larger models that go beyond literal copying, including translation and pattern matching. How far these fuzzier versions share the same mechanism is still open.

Get started: learn AI alignment theory step by step

Induction heads are a good entry point into interpretability, but they make more sense once you know how modern models are built and trained. Learn AI Alignment Theory is built for that path. It has 22 courses split into three levels. The Basic level includes How Modern AI Works, with no background needed. The Advanced level includes a full Interpretability course, alongside Scalable Oversight and The Science of LLM Misalignment.

Lessons take about 8 minutes each. You answer questions where you choose, put in order, match, sort and estimate a number. Hands-on activities include scenarios where you switch assumptions on and off. Each lesson separates what is known from what is still open and lists its sources. Where researchers disagree, debate cards state each serious position fairly with no verdict. Finished lessons come back as spaced review, so ideas like K-composition stay with you. A 349-term glossary and a topics map follow the aisafety.com self-study topics.

You sign in with Google or an emailed code. You can read more on the about page.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.