AI sandbagging is when an AI model does worse on a test than it really can, on purpose. Labs use test scores to decide whether a model is safe to release, and those scores are becoming part of AI regulation. If a model…
AI evaluations
4 posts
AI sandbagging explained: when models hide what they can do
AI jailbreaks explained: why safety training can fail
AI jailbreaks are inputs that get a safety-trained language model to do what its safety training was meant to stop. A model is trained to refuse certain requests, and a jailbreak finds a way around that refusal. Studying…
Dangerous capability evaluations: how AI labs test for risk
Dangerous capability evaluations are tests that ask one narrow question about an AI model: could it do something that would help cause severe harm? Examples are writing working cyber attacks, persuading people against…
AI red teaming: how researchers look for dangerous behavior
Before an AI model is released, someone has to try to make it misbehave. AI red teaming is that job: deliberately attacking a model with tricky, hostile or unusual inputs to find harmful behavior before real users do.…