Prompt Leaking
Prompt leaking is an attack that tricks a model into revealing its hidden system prompt or other confidential instructions.
Description
System prompts should be treated as potentially discoverable, so they should not contain secrets such as API keys.
Sources
- Perez & Ribeiro (2022). Ignore Previous Prompt: Attack Techniques For Language Models.
Cite this entry
Protologue. (2026). Prompt Leaking. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0089). https://protologue.com/t/prompt-leaking/
BibTeX
@misc{protologue_prompt_leaking,
title = {Prompt Leaking},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0089},
url = {https://protologue.com/t/prompt-leaking/}
}