protologue

Prompt Compression

Also called LLMLingua, context compression.

Prompt compression shortens a prompt by removing tokens that contribute little information, typically scored by a smaller language model, to cut cost and latency while preserving task performance.

Description

LLMLingua used a small model's perplexity to drop low-information tokens and reported high compression ratios with limited performance loss.

Sources

  1. Jiang et al. (2023). LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

Cite this entry

Protologue. (2026). Prompt Compression. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0080). https://protologue.com/t/prompt-compression/

BibTeX
@misc{protologue_prompt_compression,
  title = {Prompt Compression},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0080},
  url = {https://protologue.com/t/prompt-compression/}
}

Markdown JSON