protologue

Prompt Caching

Also called context caching, prefix caching.

Prompt caching stores the model's processed state for a reused prompt prefix, such as a long system prompt or document, so later requests sharing that prefix are cheaper and faster.

Description

Because caches match on an exact prefix, prompts are structured with stable content first and variable content last.

Sources

  1. Anthropic (2024). Prompt caching.

Cite this entry

Protologue. (2026). Prompt Caching. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0081). https://protologue.com/t/prompt-caching/

BibTeX
@misc{protologue_prompt_caching,
  title = {Prompt Caching},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0081},
  url = {https://protologue.com/t/prompt-caching/}
}

Markdown JSON