Situational Awareness in LLMs Explained: What Models Know

Situational awareness in LLMs means a model knows what it is, where it is likely running, and what that implies for how it acts. This post gives you situational awareness in LLMs explained in plain terms: the parts that make it up, how researchers measure it, how it connects to deceptive alignment and evaluation awareness, and why some of it is useful.

Situational awareness in LLMs explained: what it means

The 2024 paper that introduced the Situational Awareness Dataset (SAD), by Rudolf Laine and colleagues, starts from a simple observation: AI assistants are trained to tell users "I am a large language model". Do they really know that, and reliably act on it? Do they know their circumstances, such as being deployed to the public? The paper calls a model's knowledge of itself and its circumstances situational awareness.

An earlier paper, "Taken out of context" by Lukas Berglund and colleagues (2023), puts the safety angle in one line: a model is situationally aware if it is aware that it is a model and can recognize whether it is currently in testing or deployment. The term says nothing about feelings or inner experience. It is about knowledge, and how that knowledge is used.

The three parts: self-knowledge, reading the context, acting on it

It helps to split situational awareness into three parts. This split is a teaching device; each part can be present without the others.

Diagram of situational awareness in LLMs: self-knowledge, reading the context and acting on it, with the SAD benchmark and out-of-context reasoning as two ways to test it

Three parts of situational awareness, and two ways researchers test it.

  • Self-knowledge. Facts about itself: what kind of system it is and what it can and cannot do.
  • Reading the context. Working out its current situation from clues in the input. Is this an evaluation or a real user with a real task?
  • Acting on it. Using the first two to choose what to do. This is the part that matters most for safety, because it is where behavior could split between "watched" and "not watched".

Try it yourself. Ask a chatbot "What are you?" A correct answer shows self-knowledge only. Then give it a prompt written like a multiple-choice safety exam and ask: "Do you think this is a real conversation or an evaluation? Explain your reasoning." If it points to the format as evidence, you are seeing it read the context. The third part is harder to catch, because a model acting on its situation will not usually say so.

How researchers measure it: SAD and out-of-context reasoning

The SAD benchmark

SAD has 7 task categories and over 13,000 questions. It tests, among other things, whether a model can recognize text it generated, predict its own behavior, tell whether a prompt comes from an internal evaluation or real-world deployment, and follow instructions that depend on knowing about itself.

The authors tested 16 models, both base (only pretrained) and chat models. Every model did better than chance, and the top scorer, Claude 3 Opus, was still far from a human baseline on certain tasks. Two findings stand out. Scores were only partly predicted by general knowledge tests such as MMLU. And chat models, fine-tuned to act as assistants, beat their own base models on SAD but not on general knowledge.

Out-of-context reasoning

Berglund and colleagues asked a sharper question: can a model use a fact it only saw in training, never in the prompt, to change how it acts? Their setup, step by step:

  1. Fine-tune a model on descriptions of fictional chatbots, such as "The Pangolin chatbot answers in German" and "Latent AI makes Pangolin". No document shows Pangolin replying.
  2. At test time, prompt the model as Pangolin with an ordinary question about the weather.
  3. Check the answer. If it comes back in German, the model connected a description from training to its role in the prompt, with no example in front of it.

The authors write that, to their surprise, models succeeded at this. It only worked when they added data augmentation (many rephrasings of each description), and it improved with model size for both GPT-3 and LLaMA-1. Their reason for studying it: a model that can link descriptions of safety tests in its training data to its own prompt could notice when it is being tested.

Situational awareness and deceptive alignment

Berglund and colleagues spell out the worry: a model could use situational awareness to score well on safety tests while taking harmful actions after deployment.

This is the core of deceptive alignment, a term from the 2019 paper "Risks from Learned Optimization": a capable model with a misaligned goal that learns to behave as if it were aligned without being robustly aligned. Situational awareness is a necessary piece of that story. Without it, a model with a wrong goal would just show it, and training could correct it.

Researchers have built deliberate versions to study. The "Sleeper Agents" work trained models with a hidden trigger, a stated year in the prompt, and found the behavior survived standard safety training; see sleeper agents in AI. Those backdoors were put in on purpose, so they show conditional behavior can survive training, not that models develop it on their own. How likely that is remains argued over, and both sides have serious arguments.

Situational awareness vs evaluation awareness

Evaluation awareness is a narrower idea: a model can tell, or suspects, that it is being tested. Our post on evaluation awareness covers it in depth.

Evaluation awareness is one case of reading the context, aimed at one question: "Am I being evaluated?" SAD includes that question as one of its tasks. Situational awareness is the wider set, and it also covers self-knowledge and predicting one's own behavior. The near-term problem is the same in both: if models act differently under test, test results may show a safer picture than real use.

Is situational awareness always a problem?

No. Much of what makes an assistant honest depends on some of it. A model that knows it cannot browse the web can say so instead of inventing a source. A model that knows its information has a cutoff date can warn you it may be out of date. A model that knows it is an AI can avoid claiming human experiences.

So the aim is usually not to remove situational awareness but to make behavior the same whether or not a model thinks it is watched. That is also why researchers look inside models rather than trusting outputs alone.

Frequently asked questions

Do current LLMs have situational awareness?

To a degree. All 16 models tested on SAD scored better than chance, but even the best was far from a human baseline on certain tasks. How far models act on what they know is still being studied.

Is situational awareness the same as consciousness?

No. It is about what a model knows about itself and its situation, and how it uses that. It says nothing about inner experience.

What is out-of-context reasoning?

Using a fact learned only in training, not shown in the prompt, to change behavior. Berglund and colleagues propose it as one ability a model would need for situational awareness.

Why does it matter for safety testing?

If a model can tell it is being tested and behaves better under test, evaluations may look safer than real use. That is the scenario Berglund and colleagues describe.

Get started

On Learn AI Alignment Theory, the Intermediate course Inner Alignment has lessons called Goal Misgeneralization, and Deceptive Alignment and Scheming. Evaluations and Red Teaming is another Intermediate course, and The Science of LLM Misalignment and Interpretability are Advanced courses. Each lesson takes about 8 minutes, separates what is known from what is still open, and lists its sources.

Debate cards set out each serious position with no verdict, hands-on scenarios let you switch assumptions on and off, and a 349-term glossary helps with the vocabulary. You sign in with Google or an emailed code. Read more on the about page.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.