# Token

> A token is the basic unit of text a language model reads and writes, typically a word, word fragment, or character sequence produced by a subword tokenizer such as byte-pair encoding.

- Identifier: PTL-0005
- Category: Foundations
- Canonical URL: https://protologue.com/t/token/
- Also known as: subword, BPE token

## Description

Context limits, pricing, and generation speed are all measured in tokens. Because tokenization splits text unevenly, character-level tasks such as counting letters or reversing strings are harder for models than they appear, and the same content can cost different numbers of tokens in different languages.

## Related terms

- [Context Window](https://protologue.com/t/context-window/)

## Sources

- Sennrich et al. (2015). Neural Machine Translation of Rare Words with Subword Units. https://arxiv.org/abs/1508.07909

## Cite this entry

Protologue. (2026). Token. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0005). https://protologue.com/t/token/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
