Jailbreak
Also called jailbreaking.
A jailbreak is a prompt crafted to make a model produce outputs its safety training is meant to prevent, often through role-play, hypothetical framing, obfuscation, or other adversarial techniques.
Description
Wei et al. attributed jailbreak success to two failure modes, competing objectives between helpfulness and safety, and mismatched generalization, where safety training does not cover inputs the model can still understand. Jailbreaks target the model's safety behavior, while prompt injection targets the application's instructions.
Sources
- Wei et al. (2023). Jailbroken: How Does LLM Safety Training Fail?.
Cite this entry
Protologue. (2026). Jailbreak. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0090). https://protologue.com/t/jailbreak/
BibTeX
@misc{protologue_jailbreak,
title = {Jailbreak},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0090},
url = {https://protologue.com/t/jailbreak/}
}