protologue

Prompt Leaking

Prompt leaking is an attack that tricks a model into revealing its hidden system prompt or other confidential instructions.

Description

System prompts should be treated as potentially discoverable, so they should not contain secrets such as API keys.

Sources

  1. Perez & Ribeiro (2022). Ignore Previous Prompt: Attack Techniques For Language Models.

Cite this entry

Protologue. (2026). Prompt Leaking. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0089). https://protologue.com/t/prompt-leaking/

BibTeX
@misc{protologue_prompt_leaking,
  title = {Prompt Leaking},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0089},
  url = {https://protologue.com/t/prompt-leaking/}
}

Markdown JSON