# Prompt Compression

> Prompt compression shortens a prompt by removing tokens that contribute little information, typically scored by a smaller language model, to cut cost and latency while preserving task performance.

- Identifier: PTL-0080
- Category: Prompt Optimization
- Canonical URL: https://protologue.com/t/prompt-compression/
- Also known as: LLMLingua, context compression
- Introduced: 2023

## Description

LLMLingua used a small model's perplexity to drop low-information tokens and reported high compression ratios with limited performance loss.

## Related terms

- [Context Engineering](https://protologue.com/t/context-engineering/)
- [Prompt Caching](https://protologue.com/t/prompt-caching/)

## Sources

- Jiang et al. (2023). LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. https://arxiv.org/abs/2310.05736

## Cite this entry

Protologue. (2026). Prompt Compression. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0080). https://protologue.com/t/prompt-compression/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
