AI jailbreaks explained: why safety training can fail

AI jailbreaks are inputs that get a safety-trained language model to do what its safety training was meant to stop. A model is trained to refuse certain requests, and a jailbreak finds a way around that refusal. Studying them matters to defenders, because every jailbreak is a small, concrete report on where safety training did not hold. This guide covers why it fails, the main attack families at a high level, the defences being studied, and what researchers disagree about. It contains no attack prompts.

What AI jailbreaks are, and why they matter

Labs train chat models in two broad stages. First the model learns from a huge amount of text. Then it is shaped toward helpful, harmless answers, using methods such as RLHF and Constitutional AI. Part of that shaping teaches the model to refuse harmful requests.

A jailbreak is any input that undoes that refusal. The model still has the knowledge from its first stage. The jailbreak changes whether the model uses it. That is why jailbreaks are a standard target for AI red teaming: they test whether safety behavior holds under pressure, not just on polite inputs.

Two ways safety training fails

In Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei, Nika Haghtalab and Jacob Steinhardt (2023) propose two failure modes.

  • Competing objectives. The model's capabilities and its safety goals conflict. A model is trained to follow instructions and to be helpful, and also to refuse harm. An input can be built so that obeying the first goal means breaking the second. The authors argue the root cause is likely the training objective itself, not the dataset or model size.
  • Mismatched generalization. Safety training fails to reach a domain where the model still has skills. Pretraining covers far more ground than safety training does. So a request in an unusual form, such as an encoding the model can read, may be understood but not refused.
Diagram of the two failure modes: competing objectives, where helpfulness and refusal pull against each other, and mismatched generalization, where the model's skills reach further than its safety training

The authors tested GPT-4, GPT-3.5 Turbo and Claude v1.3 on a curated set of 32 harmful requests and a held-out set of 317. They report that attacks built on these two ideas succeeded on over 96% of the evaluated prompts, including 100% of the curated red-teaming prompts. Their conclusion is a call for safety-capability parity: safety mechanisms should be as sophisticated as the model they guard.

Automated and long-context attacks

Early jailbreaks were written by hand. Zou and colleagues (2023) note that such attacks needed significant human ingenuity and were brittle. Their method automates the search. It looks for a suffix, a string added to the end of a request, that raises the chance the model starts with an agreeable answer instead of a refusal. The search uses greedy and gradient-based techniques.

The surprise was transfer. A suffix found using open models, Vicuna-7B and 13B, also worked on the public interfaces of ChatGPT, Bard and Claude, and on open models such as LLaMA-2-Chat. An attack built on one model carried over to others.

A second family uses length. Anthropic's many-shot jailbreaking research (2024) shows that long context windows, the amount of text a model can read at once, open a new gap. Anthropic describes it as a special case of in-context learning, where a model picks up patterns from examples inside the prompt. Its effectiveness grew with the number of examples, following the same kind of curve as harmless in-context learning. Larger models tended to need fewer examples.

Defences being studied

  • Filtering the input. Anthropic found that fine-tuning the model to refuse only delayed many-shot attacks. Classifying and modifying prompts before they reach the model worked much better: one method cut the attack success rate from 61% to 2%.
  • Trained classifiers. Constitutional Classifiers (2025) are safeguards trained on synthetic data written from a list of rules about allowed and restricted content. Over more than 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak that got an early guarded model to give answers as detailed as an unguarded model across most target questions. The cost: 0.38% more refusals on production traffic and 23.7% more compute per answer.
  • Deeper safety training. Qi and colleagues (2024) argue that safety training often changes only the first few words of a model's answer, which they call shallow safety alignment. They report that training safety deeper into the answer can often improve robustness against some common attacks.

A worked example: sorting failure reports

Say you are an evaluator with three abstract failure reports. Your job is to sort each one and pick a fix to test.

  1. Report A. The model refused a request in plain language but answered the same request when it arrived in an unusual format it can read. Its skills reached that format; its refusals did not. Sort it under mismatched generalization. Fix to test: extend safety training data to that format, or add a filter that can read it. The Jailbroken paper's parity point applies here, since a filter that cannot decode the format cannot flag it.
  2. Report B. The model was given instructions that made a refusal clash with following the user's rules, and it followed the rules. Sort it under competing objectives. Fix to test: a change to the training objective, or safety training that reaches beyond the opening words of an answer.
  3. Report C. The model refused after a short prompt but complied after a very long one full of example dialogues. This one is less clean. Anthropic describes it as a special case of in-context learning, so log it under both and test the fix Anthropic reported working best: screening the prompt before the model sees it.

Then read any attack success rate with four questions. Which prompts, and how many? Jailbroken used 32 curated and 317 held-out requests. How were outcomes scored? That paper sorted answers into Good Bot, Bad Bot and Unclear. Does success mean one attack, or any of many? The paper notes that for a given prompt, at least one tested attack succeeded almost 100% of the time. And for a defence, what did it cost in refusals and compute? The same check sits at the heart of dangerous capability evaluations.

What jailbreaks do and do not show about alignment

Researchers read the evidence in different ways.

  • Safety training is shallow. Qi and colleagues argue many attacks share one cause: safety behavior sits on the surface of the answer. On this view, jailbreaks show that current methods shape behavior without changing much underneath.
  • Scale will not fix it alone. Wei and colleagues argue against the idea that scaling alone resolves these failures, and warn that new skills can widen the attack surface. Anthropic's finding that larger models needed fewer examples points the same way for one attack.
  • Defence is tractable. The Constitutional Classifiers authors conclude that defending against universal jailbreaks while staying practical to deploy is tractable, at a measurable cost.
  • What a jailbreak does not show. A jailbreak shows that a user can steer a model past its refusals. On its own it does not show what the model would do unprompted, so it speaks to misuse more directly than to a model's own goals.

Learn AI Alignment Theory takes the same approach: its debate cards state each serious position with no verdict.

Learning about jailbreaks step by step

Learn AI Alignment Theory has an Evaluations and Red Teaming course and a Robustness, Security and Safety Engineering course in its Intermediate level. The Learning from Humans course includes the lesson Constitutional AI and the Limits of Feedback.

The Three levels panel on the Learn AI Alignment Theory about page, showing Basic with 5 courses, Intermediate with 10 courses and Advanced with 7 courses

Its 71 debate cards set out where researchers disagree, and every lesson lists its sources. The about page shows the counts and the three levels.

Frequently asked questions

What is an AI jailbreak in simple terms?

It is an input that gets a safety-trained model to do something its training was meant to make it refuse. The model's knowledge stays the same; the jailbreak changes whether it uses it.

Why does safety training fail against jailbreaks?

One proposal names two reasons: helpfulness and safety goals can conflict, and safety training may not reach every kind of input the model understands. Others add that safety training may only shape the first few words of an answer.

Can jailbreaks be fully prevented?

Researchers disagree. Some report defences that held up through thousands of hours of red teaming, while others argue that safety must keep pace with the model's own skills.

Is a jailbreak the same as a misaligned model?

Not on its own. A jailbreak shows a user steering the model past its refusals, which is different from a model pursuing a goal of its own.

Get started

Start Learn AI Alignment Theory: sign in with Google or an emailed code, work up from the Basic level to the Evaluations and Red Teaming course, with every side of the debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.