Activation Steering Explained: Changing AI Behavior From Inside
Most ways of shaping a language model work from the outside: you write a better prompt, or you train it on new examples. Activation steering works from the inside. You reach into the model while it runs and nudge its internal state toward the behavior you want. This is activation steering explained: how a steering vector is built, what four research papers found, how it compares to other methods, and why it cannot by itself show a model is aligned.
Activation steering explained in plain words
Activation steering means adding a chosen direction to a model's internal numbers while it processes text, so its output shifts in a predictable way. Nothing is retrained. Alexander Matt Turner and colleagues, who introduced Activation Addition in 2023, describe the broader idea as activation engineering: changing activations at the moment the model runs in order to steer what it writes.
An analogy helps. A prompt is like asking a driver to turn left. Finetuning is like teaching the driver new habits over many lessons. Steering is like putting your hand on the wheel during one drive. The driver is the same; only this trip changes.
How activations and the residual stream work
A transformer processes text as a sequence of tokens. At each layer, every token carries a long list of numbers, its activation. These lists travel through the model along the residual stream, a shared running sum: each block reads from it, computes something, and adds its result back.
Steering rests on one observation: many concepts seem to line up with directions in this space. Move an activation along one direction and the output may become more positive; along another, it may change topic. If you want the background on how researchers find and test such directions, the linear probes and superposition posts cover it. It is a working idea that holds well in some cases and poorly in others.

How to build a steering vector
The simplest recipe uses a pair of prompts that differ in one property. Here is the routine, which you can follow on an open model with a library that lets you read and change activations:
- Write a contrast pair. Turner and colleagues use pairs such as
LoveversusHate. - Record activations. Run both through the model and save the residual stream at one layer.
- Subtract. The difference, love minus hate, is your steering vector.
- Add it while generating. On a new prompt, add the vector at the same layer, multiplied by a strength you choose.
- Sweep the strength and layer. Try several values and measure both the effect you want and any damage to the rest of the output.
Turner and colleagues report that this needs no optimization and works with a single pair, so you can try ideas quickly. They report control over topic and sentiment while keeping performance on unrelated tasks.
Nina Panickssery and colleagues made it sturdier with Contrastive Activation Addition. Instead of one pair, they average the difference over many pairs of positive and negative examples of a behavior, such as factual versus made-up answers. The vector is added at every token after the user's prompt, with a positive strength to increase the behavior or a negative one to reduce it. On Llama 2 Chat, they report it changed behavior significantly, worked on top of finetuning and system prompts, and only slightly reduced capabilities.
What steering has shown so far
Refusal sits along one direction
Andy Arditi and colleagues found that, in 13 open chat models up to 72 billion parameters, refusal was carried by a single direction. Erasing it stopped the model refusing harmful requests; adding it made the model refuse even harmless ones. They used this as a way to switch off refusal in open models and say it shows how brittle current safety finetuning is. The AI jailbreaks post covers why that matters.
Truthfulness can be nudged
Kenneth Li and colleagues shifted activations along directions in a small number of attention heads, a method they call Inference-Time Intervention. On the Alpaca model, truthfulness on the TruthfulQA benchmark rose from 32.5% to 65.1%. They also found a tradeoff: pushing truthfulness harder made the model less helpful, so the strength has to be tuned.
A wider frame: representation engineering
Andy Zou and colleagues group this kind of work under representation engineering: studying patterns across many units rather than single neurons, both to monitor and to change high-level behavior. They apply it to honesty, harmlessness and power-seeking.
Activation steering vs prompting, finetuning and RLHF
- Prompting changes the input. Cheap and easy to undo, but a long conversation or a clever user can override it.
- Finetuning changes the weights. Lasting and strong, but it needs data and can spread in surprising ways, as the emergent misalignment results show.
- RLHF changes the weights using a reward model trained on people's comparisons, and inherits that reward's flaws.
- Steering changes activations while the model runs. The weights stay the same, and you can switch it on, off or scale it per request.
Steering does not replace the others. It is fast to try, easy to reverse, and gives a way to ask what the model represents inside.
Why it matters for alignment, and its limits
Training only sees behavior, while what matters is the goal producing it. Steering offers one way to look past outputs: if removing a direction changes behavior, that is evidence the model was using it. Li and colleagues read their result as a hint that models may hold an internal sense of whether something is true even while saying something false.
The limits are real:
- Side effects. A vector rarely captures one clean concept, and too much strength can break the output.
- Tradeoffs. Turning one behavior up can turn another down, as with truthfulness and helpfulness.
- What you contrast is what you get. If your "honest" examples are also shorter, you may have built a "short" vector.
- Dual use. The method that adds refusal can remove it from an open model.
- It does not change what was learned. Steering adjusts behavior while the model runs; it does not show that the model's goals are what you want.
Researchers differ on how much weight steering results deserve. Some see them as evidence that high-level concepts are simple directions we can use for monitoring and control. Others stress the side effects and the gap between changing an output and understanding the model. The mechanistic interpretability for beginners post sets steering beside the bottom-up approach.
Frequently asked questions
Is activation steering the same as representation engineering?
They overlap. Representation engineering, as Zou and colleagues frame it, covers both monitoring and changing high-level representations; activation steering usually means the changing part.
Does activation steering change the model's weights?
No. You add a vector to activations while the model runs, and the weights stay as they were. Turn the vector off and the model behaves as before.
Can steering make a model more truthful?
In one study, yes on a benchmark: Li and colleagues raised Alpaca's TruthfulQA score from 32.5% to 65.1%, with a cost to helpfulness. That does not show the model became honest in any deep sense.
Which layer should you steer at?
It depends on the model and the behavior. Sweep several layers and strengths, and keep the setting that gives the effect with the fewest side effects.
Get started
Steering builds on ideas you need first: what a model computes, how goals get specified, and how training can produce aims nobody intended. Learn AI Alignment Theory covers these in 22 courses and 69 lessons of about 8 minutes. The Basic level starts with How Modern AI Works and Specifying Goals. The Intermediate Inner Alignment course has Goal Misgeneralization and Deceptive Alignment and Scheming. The Advanced level has Interpretability and The Science of LLM Misalignment.
Every lesson lists its sources and separates what is known from what is still open, and debate cards state each serious position with no verdict. You sign in with Google or an emailed code.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.