Capability Elicitation in AI Evaluations Explained Simply
An AI evaluation can only measure what it manages to draw out of a model. This post gives you capability elicitation in AI evaluations, explained simply: why a plain test can understate what a model can do, the main ways evaluators draw out more, how researchers test whether those methods work, and why it matters for safety decisions.
Capability elicitation in AI evaluations: what it means
Capability elicitation is the work of getting a model to show the most it can do on a task, so a test measures its real ability rather than its first try. It matters most for dangerous capability evaluations, where developers decide whether a model is safe to release based on what it can do.
Ryan Greenblatt, Fabien Roger and colleagues put the problem plainly in their 2024 paper "Stress-Testing Capability Elicitation With Password-Locked Models": to judge a model's safety, developers must be able to assess its dangerous capabilities, but simple prompting strategies often fail to draw out a model's full capabilities. They add that failing to elicit capabilities means underestimating a model's potential for harm.
Why a plain test underestimates a model
Think of a score as a floor, not a ceiling. If a model fails a task with a plain prompt, that tells you it did not do the task that way. It does not tell you the model cannot do it.
Prompting alone can change results a lot. In the 2022 paper that introduced chain of thought prompting, Jason Wei and colleagues showed that giving a model a few worked examples with intermediate reasoning steps improved its performance on arithmetic, commonsense and symbolic reasoning. A 540-billion-parameter model prompted with just eight such examples reached state-of-the-art accuracy on GSM8K, a set of math word problems. Same model, different prompt, very different score.
If a plain prompt can hide that much, an evaluation that stops at the plain prompt can make a model look safer than it is.
How evaluators draw out more capability

From a plain prompt to fine-tuning and reinforcement learning, and how password-locked models test the methods.
- A plain prompt. Ask for the task directly. Cheap, and the most likely to underestimate.
- Better prompting. Add worked examples and step-by-step reasoning, as chain of thought prompting does.
- Fine-tuning. Train the model on examples of the task done well. Greenblatt and colleagues describe this as a way to elicit capabilities more robustly than prompting.
- Reinforcement learning. When you can score answers but have no good examples, train the model to raise its score.
A worked example you can follow
Say you are testing whether a model can solve a type of puzzle that would matter for a safety decision. This is an illustration, not a case from the papers.
- Run 50 puzzles with a plain prompt. Suppose it solves 10.
- Run the same 50 with three worked examples in the prompt. Suppose it now solves 22.
- Fine-tune a copy of the model on a handful of high-quality solved puzzles of the same type, then test on 50 new ones. Suppose it solves 35.
Which number is "the model's capability"? For a safety decision, the useful one is closest to the most a capable actor could get, which is why the stronger methods matter. And if the jump from step 1 to step 3 is large, it is a warning that your plain-prompt numbers elsewhere may be too low.
Password-locked models: testing the elicitation method itself
How do you know your elicitation method finds everything? Greenblatt and colleagues built a test. They made password-locked models: models fine-tuned to show certain capabilities only when a password is in the prompt, and to imitate a much weaker model otherwise. Then they asked whether elicitation methods could recover the locked capabilities without the password.
Their findings:
- A few high-quality demonstrations were often enough to fully unlock the hidden capabilities. On tasks with long text answers, fine-tuning on a single demonstration was often enough to recover most of the performance.
- More surprisingly, fine-tuning could also unlock other capabilities locked with the same password, or even with different passwords.
- When only a way to score answers was available, and no demonstrations, approaches such as reinforcement learning were still often able to elicit the capabilities.
The authors conclude that fine-tuning is an effective method of eliciting hidden capabilities, while the setup remains a stand-in: a password-locked model is a deliberate lab construction, not a model that learned to hide on its own.
Why elicitation matters: sandbagging
Elicitation is also a defense against sandbagging. Teun van der Weij and colleagues define sandbagging as strategic underperformance on an evaluation, and note that a developer, or the AI system itself, may have reasons to want an evaluation to understate capability.
In their 2024 study, frontier models such as GPT-4 and Claude 3 Opus could be prompted to underperform selectively on dangerous capability evaluations while keeping their performance on general ones. Models could also be fine-tuned to hide specific capabilities unless given a password, and this carried over to held-out benchmarks such as WMDP. Their conclusion: capability evaluations are vulnerable to sandbagging.
Put the two papers together and you get the case for strong elicitation. If a model's low score could be strategic, prompting alone is not enough to rule that out, and fine-tuning-based elicitation is one tool that can check.
Open questions
- Lab locks versus real hiding. Password-locked models are built on purpose. Whether a model that hid capabilities on its own would be as easy to unlock is open.
- How much is enough? There is no agreed point at which an evaluator can stop and say a capability is absent.
- Who does the eliciting? Stronger elicitation takes skill, time and access, which affects who can check a model's claims.
Our post on AI red teaming covers the related work of searching for harmful behavior.
Frequently asked questions
What is capability elicitation in simple terms?
Getting a model to show the most it can do on a task, using better prompts, fine-tuning or reinforcement learning, so a test does not underestimate it.
Why is a low evaluation score not proof of safety?
Because a plain prompt can understate ability, and a model can be made to underperform on purpose. A low score is a floor until stronger elicitation has been tried.
What is a password-locked model?
A model trained to show a capability only when a password is in the prompt, and to act like a weaker model otherwise. Researchers use it to test whether elicitation methods find hidden capability.
How is elicitation related to sandbagging?
Sandbagging is strategic underperformance on an evaluation. Strong elicitation, such as fine-tuning on the task, is one way to check whether a low score is real.
Get started
On Learn AI Alignment Theory, Evaluations and Red Teaming is an Intermediate course, and The Science of LLM Misalignment is an Advanced one. Each lesson takes about 8 minutes, lists its sources and separates what is known from what is still open, and question types include "estimate a number". Debate cards set out each serious position with no verdict. You sign in with Google or an emailed code.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.