# Jailbreak

> A jailbreak is a prompt crafted to make a model produce outputs its safety training is meant to prevent, often through role-play, hypothetical framing, obfuscation, or other adversarial techniques.

- Identifier: PTL-0090
- Category: Security & Adversarial Prompting
- Canonical URL: https://protologue.com/t/jailbreak/
- Also known as: jailbreaking

## Description

Wei et al. attributed jailbreak success to two failure modes, competing objectives between helpfulness and safety, and mismatched generalization, where safety training does not cover inputs the model can still understand. Jailbreaks target the model's safety behavior, while prompt injection targets the application's instructions.

## Narrower terms

- [Adversarial Suffix](https://protologue.com/t/adversarial-suffix/)
- [Many-shot Jailbreaking](https://protologue.com/t/many-shot-jailbreaking/)

## Related terms

- [Prompt Injection](https://protologue.com/t/prompt-injection/)
- [Constitutional AI](https://protologue.com/t/constitutional-ai/)

## Sources

- Wei et al. (2023). Jailbroken: How Does LLM Safety Training Fail?. https://arxiv.org/abs/2307.02483

## Cite this entry

Protologue. (2026). Jailbreak. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0090). https://protologue.com/t/jailbreak/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
