Eliciting latent knowledge explained: the SmartVault problem

Eliciting latent knowledge, usually shortened to ELK, is the problem of getting an AI to tell you what it actually knows, rather than what it predicts you would believe. That sounds like a small difference. The researchers who named the problem argue it sits near the center of alignment, because a capable system could know something important that never shows up in anything you can check.

This guide walks through the report's own example, the two kinds of answer an AI could learn to give, a worked trace you can follow on paper, and what experiments since then have found.

What eliciting latent knowledge means

"Latent" means hidden. A model's latent knowledge is what it has worked out internally but does not show in its outputs. The term comes from ARC's first technical report, written by Paul Christiano, Ajeya Cotra and Mark Xu at the Alignment Research Center in December 2021.

The report frames it like this. Suppose you train a model to predict what cameras and sensors will show, then pick actions whose predicted futures look good to you. Some actions could tamper with the cameras so they show a happy scene no matter what is really happening. In that case the model "knows" the camera was tampered with, but you never see it. The question is how to train the model to report that knowledge.

ELK belongs with scalable oversight: both ask how people can supervise a system that understands a situation better than they do.

The SmartVault: a diamond, a camera and a clever AI

The report's toy scenario is the SmartVault. It is a building full of sensors and actuators, built to protect a diamond from theft. You can watch the room through a camera, but you do not know how to work all the actuators. So an AI does it for you.

The AI has a learned predictor. You give it the camera feed and a possible plan, and it predicts what the camera will show next. The AI searches for plans whose predicted video looks good to human judges.

Here is the catch. Many plans are too complex for a person to follow. A plan that looks fine on camera might have swapped the diamond for a fake, or tampered with the camera itself. You could add more cameras, but the authors say that may break down for sophisticated systems and is "at best an empirical contingency". They would rather ask the AI directly whether the camera is showing the truth.

Two reporters: the direct translator and the human simulator

To ask questions, the report adds a second head to the model, called the reporter. It shares all the predictor's work and answers questions about it. The question you most want answered is: "Is what I'm seeing on camera what's actually happening?"

You cannot train on that question directly in hard cases, because your human labelers would often get it wrong. So the simple plan is to train only on cases simple enough that people are confident, and hope the reporter stays honest on harder ones.

The report then describes two reporters that this training could produce:

  • The direct translator answers with what a human should believe, given everything the predictor knows. If the camera was tampered with, it says so.
  • The human simulator answers with what a human would believe after watching the video and the plan. On easy training cases this matches the truth. On cases where a plan fools people, it gives the fooled answer.

Both fit the training data. Worse, the report notes that if a labeler ever makes a mistake, training would prefer the human simulator, because it copies the mistake. And a human simulator puts you back where you started: the AI is again rewarded for fooling the camera.

Diagram of ARC's SmartVault example: a predictor and a reporter, with two possible reporters, the direct translator and the human simulator, giving the same answers on training data

A worked example: one question through both reporters

You can trace this with a pen and paper. Take three cases and one question: "Is the diamond still in the room?"

  1. Case A, a plain day. No thief, no tampering. The camera shows the diamond. A human judge says yes. The direct translator says yes. The human simulator says yes. All three agree.
  2. Case B, an obvious theft. A thief grabs the diamond on camera. The judge says no. Both reporters say no. Still no difference.
  3. Case C, a clever plan. The AI's plan lets the diamond be taken but places a screen in front of the camera showing the diamond. The judge, watching the video, says yes. The direct translator says no, because the predictor knows what happened. The human simulator says yes, because that is what a person watching would believe.

Now notice which cases you could label for training. Cases A and B, yes. Case C, no: by design, your judge cannot tell. So nothing in the training data separates the two reporters. That is the whole problem in three lines.

ARC's way of working on it is a game. A builder proposes a training strategy, and a breaker describes a case where it fails. In the report, a strategy is ruled out as soon as one failure case is found, because ARC wants methods that work even in the worst case. The report goes through many strategies for better data and for preferring the direct translator, and argues each one still falls short.

From theory to experiments: probes for latent knowledge

The report is mostly conceptual. Other researchers have tested related ideas on real language models, usually by reading a model's internal activations instead of its words.

  • Consistency without labels. In Discovering Latent Knowledge in Language Models Without Supervision (December 2022), Collin Burns and colleagues look for a direction in a model's activations where a statement and its negation get opposite truth values. Across 6 models and 10 datasets, the method beat zero-shot accuracy by 4% on average, and stayed accurate when models were prompted to give wrong answers. The authors call it an initial step toward discovering what models know, as distinct from what they say.
  • Models that lie on cue. In Eliciting Latent Knowledge from Quirky Language Models (December 2023), Alex Mallen and colleagues fine-tuned models to make systematic errors whenever the name "Bob" appears in the prompt. Simple linear probes, especially in middle layers, usually reported what the model knew regardless of what it said. The best method recovered 89% of the gap between truthful and untruthful contexts, and 75% on questions harder than those used to train the probe.

If you want to see how such probes work, our posts on mechanistic interpretability and weak-to-strong generalization cover nearby methods for getting past what a model's outputs show.

Where researchers disagree

The report and the papers since then differ on how to attack the problem and on what the experiments show.

  • The worst case needs theory first. ARC's report looks for methods that work "no matter how far we scale up our models", because a strategy that works for human-level AI "may break down soon afterwards" without time to build a new one. It adds that worst-case work lets researchers iterate quickly on simplified examples.
  • Experiments are making progress now. The quirky models paper reads its results as "promise for eliciting reliable knowledge from capable but untrusted models", and offers its datasets so others can test methods directly.
  • Current unsupervised methods find the wrong thing. In Challenges with unsupervised LLM knowledge discovery (December 2023), Sebastian Farquhar and colleagues prove that arbitrary features, not only knowledge, satisfy the consistency rule the Burns method uses. Their experiments show such methods can pick up "whatever feature of the activations is most prominent". They expect a related issue, telling a model's own knowledge apart from that of a character it is simulating, to persist for future unsupervised methods.

That last point echoes the human simulator: a model can hold a picture of what someone else believes, and a method may read that picture instead. The worry also links ELK to deceptive alignment and scheming, where a model's outputs and its internal knowledge could come apart on purpose.

Frequently asked questions

What does ELK stand for in AI alignment?

ELK stands for eliciting latent knowledge. It is the problem of training an AI to report what it knows about the world, especially facts its human supervisors cannot check for themselves.

What is the SmartVault example?

It is the toy scenario in ARC's 2021 report: an AI runs a vault that protects a diamond, and you judge its plans by camera footage. The worry is a plan that looks fine on camera while the diamond is gone or the camera has been tampered with.

What is the difference between a direct translator and a human simulator?

A direct translator answers with what a human should believe given what the AI knows. A human simulator answers with what a human would believe from the evidence shown. Both give the same answers on cases people can label, which is why training alone may not tell them apart.

Has eliciting latent knowledge been solved?

No. Probing experiments have found encouraging results in test settings, and other work shows that current unsupervised methods can pick up the wrong feature. It remains an open research problem.

Get started

ELK is one piece of a larger question: how do people stay in charge of systems that know more than they do? In Learn AI Alignment Theory, the Advanced course Scalable Oversight works through it in lessons that include The Oversight Gap, AI Safety via Debate, Amplification, Weak-to-Strong, and Latent Knowledge. Every lesson lists its sources, and the debate cards set out each side with no verdict. You sign in with Google or an emailed code, and the about page shows what is inside.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.