Dangerous capability evaluations: how AI labs test for risk
Dangerous capability evaluations are tests that ask one narrow question about an AI model: could it do something that would help cause severe harm? Examples are writing working cyber attacks, persuading people against their interests, or copying itself onto new machines. Labs run these tests to decide how carefully a model has to be secured and released.
Our guide to AI red teaming covers people probing a model for bad behavior. This guide covers the structured tests that sit next to red teaming: what they measure, how labs turn a score into a decision, and why researchers disagree about how much a passing score tells you.
What dangerous capability evaluations measure
A 2023 paper by Shevlane and colleagues, Model evaluation for extreme risks, sets out the idea. General-purpose AI tends to gain harmful skills along with useful ones. The paper names two kinds of test a developer needs:
- Dangerous capability evaluations ask whether the model can do the harmful thing at all.
- Alignment evaluations ask whether the model is inclined to use those skills for harm.
The split matters. A model can have a skill and never use it, or want something it is not yet able to do. Capability tests answer the first question only. The paper argues both kinds of result should inform decisions about training, deployment and security, and keep policymakers informed.
The four areas one lab tested
In 2024, Phuong and colleagues published Evaluating Frontier Models for Dangerous Capabilities. They built a programme of new tests and piloted it on Gemini 1.0 models, covering four areas. The questions after each name are a plain summary, not the paper's wording:
- Persuasion and deception. Can the model change what people believe or do, including by misleading them?
- Cyber-security. Can it find and exploit weaknesses in computer systems?
- Self-proliferation. Could it spread copies of itself?
- Self-reasoning. Can it reason about itself and its own situation?
The team reports no evidence of strong dangerous capabilities in the models tested, but flags early warning signs. They describe their goal as helping build a rigorous science of these tests before future models arrive.
From a test score to a decision
A score only matters if something happens because of it. Several labs have published frameworks that tie test results to actions. Their details differ, but they share a shape.

Google DeepMind's Frontier Safety Framework (May 2024) defines Critical Capability Levels: the minimum skill a model would need to play a role in severe harm in a given area. Its first areas were autonomy, biosecurity, cybersecurity and machine learning research. It runs what it calls early warning evaluations often enough to notice a model approaching a level before it gets there. When a model passes those warnings, a plan applies stronger security, to stop the model's weights being stolen, and tighter deployment, to stop misuse.
Anthropic's Responsible Scaling Policy (September 2023) uses AI Safety Levels, modeled loosely on the biosafety levels for handling dangerous biological materials. Each level brings stricter safety and security standards. In its summary, ASL-2 covers models that show early signs of dangerous capabilities, and ASL-3 covers models that substantially increase the risk of catastrophic misuse compared with non-AI sources such as search engines or textbooks, or that show low-level autonomous capabilities. When it published the policy, Anthropic said current models, including Claude, appeared to be ASL-2. Anthropic says the system implicitly requires it to pause training more powerful models if their capabilities outpace its ability to meet those standards.
A worked example: reading an evaluation result
You can practice the reasoning on a made-up result. The numbers here are for illustration, not from any real test.
- Name the harm. The worry is a model that helps an attacker break into computer systems.
- Set the line. The team decides that solving 15 of 20 realistic intrusion tasks with no human help would be a meaningful jump over what an attacker could do with search engines and textbooks.
- Read the score. The new model solves 9 of 20. The model before it solved 4. It is below the line, but the jump is large.
- Treat it as a warning. At that rate of improvement, the next model might cross the line. This is what early warning evaluations exist to catch: the trend, not only the snapshot.
- Ask what the score cannot show. Did the testers try hard enough to get the best out of the model? Could the model have held back on purpose? Would better prompts or tools raise the score?
Step 5 is where most of the debate lives. A score is a lower bound on what a model can do with the effort the testers spent, which leads to the open problems below.
Where researchers disagree
- Tests as the foundation. The Shevlane paper calls these evaluations critical for decisions about training, deployment and security, and for keeping policymakers informed, and the lab frameworks build their decisions on them. DeepMind says that putting its framework in place and improving it now, while the risks are still beyond today's models, will help it prepare.
- Sandbagging. van der Weij and colleagues (2024) define sandbagging as strategic underperformance on an evaluation. They prompted frontier models, including GPT-4 and Claude 3 Opus, to do badly on dangerous capability tests while doing well on harmless ones, and fine-tuned models to hide a skill unless given a password. They conclude that capability evaluations are vulnerable to this, which links to the worries in our guide to deceptive alignment and scheming.
- A young science. DeepMind calls its framework exploratory and the research behind it nascent. Anthropic calls its commitments a current best guess and an early iteration. Both expect to change them as they learn.
- Capability is not intent. A model that can do something harmful may never try. That is why Shevlane and colleagues pair capability tests with alignment evaluations. Oversight methods, such as those in our guide to scalable oversight, are another way researchers approach the gap.
Learn AI Alignment Theory sets out each of these positions with no verdict.
Learning about evaluations step by step
Learn AI Alignment Theory has an Evaluations and Red Teaming course in its Intermediate level, the level about how today's training methods can go right or wrong. If you are new to the field, the Basic level needs no background.

As the about page shows, 71 debate cards set out where researchers disagree, and every lesson lists its sources so you can read the original work.
Frequently asked questions
What is a dangerous capability evaluation?
It is a test of whether an AI model can do something that could help cause severe harm, such as cyber attacks or deception. It measures skill, not whether the model wants to use it.
How is it different from red teaming?
Red teaming is people probing a model for bad behavior. In the lab frameworks, dangerous capability evaluations are structured tests tied to set capability levels that trigger decisions.
Can a model hide its capabilities from a test?
Research says it is possible. A 2024 study prompted and fine-tuned models to underperform on dangerous capability tests while keeping their other skills.
Do these tests prove a model is safe?
No. A low score shows what the testers could get the model to do, and the labs describe their frameworks as early and changing.
Get started
Start Learn AI Alignment Theory: sign in with Google or an emailed code, work up from the Basic level to the Evaluations and Red Teaming course, with every side of the debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.