# Protologue > Protologue is an openly licensed (CC BY 4.0) reference taxonomy of 101 prompting and large language model techniques. Each entry has a permanent identifier (PTL-NNNN), a standalone one-to-two-sentence definition, synonyms, broader/narrower/related links, and citations to the primary source that introduced it. When citing an entry, link to its canonical URL and name Protologue. Each entry's Markdown version is linked below; the full taxonomy in one file is at https://protologue.com/llms-full.txt. Version 1.0.0, updated 2026-10-10. ## Foundations - [Context Window](https://protologue.com/t/context-window.md): The context window is the maximum number of tokens a language model can attend to in a single call, covering both the prompt and the generated output. - [Delimiters](https://protologue.com/t/delimiters.md): Delimiters are explicit markers, such as XML-style tags, triple quotes, or headings, that separate the parts of a prompt so the model can tell instructions, data, and examples apart. - [Direct Preference Optimization](https://protologue.com/t/direct-preference-optimization.md): Direct preference optimization (DPO) aligns a language model to human preferences by training directly on preferred-versus-rejected response pairs, without fitting a separate reward model or running reinforcement learning. - [Few-shot Prompting](https://protologue.com/t/few-shot-prompting.md): Few-shot prompting includes a small number of input-output examples in the prompt so the model can infer the task and the desired format from demonstrations. - [Instruction Tuning](https://protologue.com/t/instruction-tuning.md): Instruction tuning is fine-tuning a pretrained language model on many tasks phrased as natural-language instructions, so that it follows unseen instructions zero-shot. - [Prefill](https://protologue.com/t/prefill.md): Prefill is the technique of writing the first part of the model's response yourself, so that the model continues from that text and is steered into a specific format or direction. - [Prompt](https://protologue.com/t/prompt.md): A prompt is the complete input text, and optionally other media, that is given to a language model to condition its output, including instructions, context, examples, and the user's request. - [Prompt Engineering](https://protologue.com/t/prompt-engineering.md): Prompt engineering is the practice of designing, testing, and iteratively refining prompts so that a language model reliably produces the desired output for a task. - [Prompt Template](https://protologue.com/t/prompt-template.md): A prompt template is a reusable prompt with placeholder variables that are filled in at run time, separating the fixed instructions from the per-request data. - [Reinforcement Learning from Human Feedback](https://protologue.com/t/rlhf.md): Reinforcement learning from human feedback (RLHF) trains a language model to produce outputs people prefer, by learning a reward model from human comparisons and optimizing the model against it. - [Role Prompting](https://protologue.com/t/role-prompting.md): Role prompting assigns the model a persona or professional role, such as "You are an experienced tax accountant," to shape its tone, vocabulary, and focus. - [System Prompt](https://protologue.com/t/system-prompt.md): A system prompt is a privileged instruction block, placed before the conversation, that sets a model's role, rules, tone, and constraints for every subsequent turn. - [Temperature](https://protologue.com/t/temperature.md): Temperature is a sampling parameter that rescales a model's output probabilities before a token is chosen; lower values make outputs more deterministic and higher values make them more varied. - [Token](https://protologue.com/t/token.md): A token is the basic unit of text a language model reads and writes, typically a word, word fragment, or character sequence produced by a subword tokenizer such as byte-pair encoding. - [Top-p Sampling](https://protologue.com/t/top-p-sampling.md): Top-p sampling, also called nucleus sampling, draws each next token only from the smallest set of candidates whose cumulative probability exceeds a threshold p. - [Zero-shot Prompting](https://protologue.com/t/zero-shot-prompting.md): Zero-shot prompting asks a model to perform a task from an instruction alone, without any worked examples in the prompt. ## Exemplars & In-Context Learning - [Active Prompting](https://protologue.com/t/active-prompting.md): Active prompting selects which questions to annotate with chain-of-thought exemplars by choosing those on which the model is most uncertain, measured by disagreement across sampled answers. - [Demonstration Label Sensitivity](https://protologue.com/t/demonstration-label-sensitivity.md): Demonstration label sensitivity refers to how much a model's few-shot performance depends on whether the example labels are correct; research found that randomly replacing labels often hurts performance only slightly. - [Exemplar Ordering](https://protologue.com/t/exemplar-ordering.md): Exemplar ordering is the arrangement of demonstrations within a few-shot prompt, which can swing accuracy from near state-of-the-art to near chance for the same set of examples. - [Exemplar Selection](https://protologue.com/t/exemplar-selection.md): Exemplar selection is the choice of which demonstrations to include in a few-shot prompt, commonly by retrieving the examples most semantically similar to the current input. - [Few-shot Calibration](https://protologue.com/t/few-shot-calibration.md): Few-shot calibration corrects a model's systematic biases toward particular answers, such as the most frequent or most recent label in the examples, by adjusting output probabilities measured on a content-free input. - [In-Context Learning](https://protologue.com/t/in-context-learning.md): In-context learning (ICL) is a language model's ability to perform a task by conditioning on instructions or demonstrations in its prompt, without any update to its weights. - [Many-shot In-Context Learning](https://protologue.com/t/many-shot-in-context-learning.md): Many-shot in-context learning places hundreds or thousands of demonstrations in a long-context prompt, often yielding large gains over few-shot prompting. ## Reasoning Elicitation - [Analogical Prompting](https://protologue.com/t/analogical-prompting.md): Analogical prompting asks the model to recall or generate relevant example problems and their solutions on its own before solving the target problem, removing the need for hand-written exemplars. - [Automatic Chain-of-Thought](https://protologue.com/t/auto-cot.md): Automatic chain-of-thought (Auto-CoT) builds chain-of-thought demonstrations without manual writing, by clustering questions for diversity and generating a reasoning chain for a representative of each cluster with zero-shot CoT. - [Chain-of-Thought Prompting](https://protologue.com/t/chain-of-thought.md): Chain-of-thought (CoT) prompting elicits a sequence of intermediate reasoning steps from a language model before its final answer, which improves performance on multi-step arithmetic, commonsense, and symbolic reasoning tasks. - [Contrastive Chain-of-Thought](https://protologue.com/t/contrastive-chain-of-thought.md): Contrastive chain-of-thought adds both valid and deliberately invalid reasoning demonstrations to a prompt, so the model learns which mistakes to avoid as well as what correct reasoning looks like. - [Generated Knowledge Prompting](https://protologue.com/t/generated-knowledge-prompting.md): Generated knowledge prompting first asks the model to produce relevant facts about a question, then supplies those generated facts as context when answering it. - [Graph of Thoughts](https://protologue.com/t/graph-of-thoughts.md): Graph of Thoughts (GoT) models a language model's reasoning as an arbitrary graph, in which thoughts can be combined, refined, and looped back on, generalizing chain and tree structures. - [Least-to-Most Prompting](https://protologue.com/t/least-to-most-prompting.md): Least-to-most prompting first asks the model to break a complex problem into simpler subproblems, then solves them in order, feeding each answer into the next. - [Maieutic Prompting](https://protologue.com/t/maieutic-prompting.md): Maieutic prompting generates a tree of recursive explanations for and against an answer, then infers the most logically consistent answer from the relations among them. - [Plan-and-Solve Prompting](https://protologue.com/t/plan-and-solve-prompting.md): Plan-and-solve prompting is a zero-shot method that instructs the model to first devise a plan dividing the task into subtasks and then carry out the plan step by step. - [Program of Thoughts](https://protologue.com/t/program-of-thoughts.md): Program of Thoughts (PoT) prompting expresses numerical reasoning as executable code, separating computation, done by an interpreter, from reasoning, done by the model. - [Program-Aided Language Models](https://protologue.com/t/program-aided-language-models.md): Program-aided language models (PAL) have the model write a program, typically Python, that expresses its reasoning, and then delegate the actual computation to an interpreter. - [Rephrase and Respond](https://protologue.com/t/rephrase-and-respond.md): Rephrase and Respond (RaR) asks the model to rephrase and expand the user's question in its own words before answering, reducing misunderstandings caused by ambiguous phrasing. - [Scratchpad](https://protologue.com/t/scratchpad.md): A scratchpad is a region of model output reserved for intermediate computation, which the model writes before its final answer so that multi-step calculations can be carried out explicitly. - [Self-Ask](https://protologue.com/t/self-ask.md): Self-ask prompting has the model explicitly pose and answer follow-up sub-questions before answering a multi-hop question, a format that can plug a search engine in to answer each sub-question. - [Self-Consistency](https://protologue.com/t/self-consistency.md): Self-consistency samples multiple chain-of-thought reasoning paths for the same question and returns the answer that appears most often, rather than relying on a single greedy decode. - [Self-Discover](https://protologue.com/t/self-discover.md): Self-Discover has the model compose a task-specific reasoning structure by selecting, adapting, and combining general reasoning modules, such as critical thinking or step-by-step analysis, before solving instances of the task. - [Skeleton-of-Thought](https://protologue.com/t/skeleton-of-thought.md): Skeleton-of-thought first asks the model for a brief outline of its answer, then expands each outline point in parallel, reducing end-to-end generation latency. - [Step-Back Prompting](https://protologue.com/t/step-back-prompting.md): Step-back prompting has the model first answer a more general, abstract question about the underlying principle, then use that answer to reason about the original specific question. - [System 2 Attention](https://protologue.com/t/system-2-attention.md): System 2 Attention (S2A) first prompts the model to rewrite the input so that it keeps only the relevant, unbiased content, then answers using the rewritten context. - [Thread of Thought](https://protologue.com/t/thread-of-thought.md): Thread of Thought is a prompting strategy for long, chaotic contexts that asks the model to walk through the context in manageable parts, summarizing and analyzing each before answering. - [Tree of Thoughts](https://protologue.com/t/tree-of-thoughts.md): Tree of Thoughts (ToT) lets a model explore multiple reasoning branches as a search tree, evaluating partial solutions and backtracking, rather than committing to a single left-to-right chain of thought. - [Universal Self-Consistency](https://protologue.com/t/universal-self-consistency.md): Universal self-consistency extends self-consistency to free-form outputs by asking the model itself to select the most consistent response among several samples, instead of counting exact-match answers. - [Zero-shot Chain-of-Thought](https://protologue.com/t/zero-shot-chain-of-thought.md): Zero-shot chain-of-thought prompting triggers step-by-step reasoning without examples by appending a cue such as "Let's think step by step" to the question. ## Self-Critique & Verification - [Best-of-N Sampling](https://protologue.com/t/best-of-n-sampling.md): Best-of-N sampling generates N candidate outputs and returns the one ranked highest by a verifier, reward model, or scoring function. - [Chain-of-Verification](https://protologue.com/t/chain-of-verification.md): Chain-of-Verification (CoVe) reduces hallucination by having the model draft an answer, plan verification questions about its claims, answer those questions independently, and then produce a corrected final answer. - [LLM-as-a-Judge](https://protologue.com/t/llm-as-a-judge.md): LLM-as-a-judge is the use of a strong language model to grade, score, or compare the outputs of models against criteria, as a scalable substitute for human evaluation. - [Mixture-of-Agents](https://protologue.com/t/mixture-of-agents.md): Mixture-of-Agents (MoA) arranges language models in layers, where each model receives all outputs from the previous layer as auxiliary input and an aggregator synthesizes a final response. - [Multi-Agent Debate](https://protologue.com/t/multi-agent-debate.md): Multi-agent debate has several model instances propose answers, read each other's reasoning, and revise their answers over multiple rounds until they converge. - [Process Reward Model](https://protologue.com/t/process-reward-model.md): A process reward model (PRM) scores each intermediate step of a model's reasoning, rather than only the final answer, and is used to select or train better reasoning chains. - [Reflexion](https://protologue.com/t/reflexion.md): Reflexion is an agent technique in which, after a failed attempt, the model writes a verbal reflection on what went wrong and stores it in memory to guide its next attempt. - [Self-Refine](https://protologue.com/t/self-refine.md): Self-Refine is an iterative method in which the same model generates an output, critiques it with specific feedback, and revises it, repeating until a stopping condition is met. ## Retrieval & Tool Use - [Function Calling](https://protologue.com/t/function-calling.md): Function calling, or tool use, is a model capability in which the model outputs a structured request to invoke a developer-defined function with arguments, which the application executes and returns as a result to the model. - [Hypothetical Document Embeddings](https://protologue.com/t/hypothetical-document-embeddings.md): Hypothetical Document Embeddings (HyDE) improves retrieval by having a model write a hypothetical answer to the query, embedding that answer, and searching for real documents similar to it. - [Model Context Protocol](https://protologue.com/t/model-context-protocol.md): The Model Context Protocol (MCP) is an open standard, introduced by Anthropic in 2024, that defines how applications expose tools, data resources, and prompt templates to language-model clients through a common client-server interface. - [ReAct](https://protologue.com/t/react.md): ReAct is a prompting pattern that interleaves reasoning traces ("Thought") with actions such as tool calls ("Action") and their results ("Observation"), letting a model plan, act, and update its plan in a loop. - [Retrieval-Augmented Generation](https://protologue.com/t/retrieval-augmented-generation.md): Retrieval-augmented generation (RAG) supplies a language model with passages retrieved from an external corpus at query time, so its output is grounded in that information rather than only in its trained parameters. - [Self-RAG](https://protologue.com/t/self-rag.md): Self-RAG trains a model to decide when to retrieve, and to emit special reflection tokens that critique whether retrieved passages are relevant and whether its own output is supported by them. - [Structured Outputs](https://protologue.com/t/structured-outputs.md): Structured outputs constrain a language model to produce responses that conform to a specified format, typically a JSON Schema, either through instructions or through constrained decoding that guarantees validity. - [Toolformer](https://protologue.com/t/toolformer.md): Toolformer is a method in which a language model teaches itself to use external tools, such as a calculator or search API, by generating candidate API calls in text and keeping those that reduce its prediction loss. ## Agents & Orchestration - [Agent Memory](https://protologue.com/t/agent-memory.md): Agent memory is the set of mechanisms that let a language-model agent store information beyond a single context window, such as conversation summaries, retrievable records of past events, and reflections, and bring the relevant parts back into context later. - [Agentic Workflow](https://protologue.com/t/agentic-workflow.md): An agentic workflow is a system in which language models and tools are orchestrated through predefined code paths, as opposed to an agent that chooses its own steps. - [AI Agent](https://protologue.com/t/ai-agent.md): An AI agent is a system in which a language model dynamically directs its own process and tool use in a loop, deciding what actions to take based on environment feedback until a task is complete. - [Context Engineering](https://protologue.com/t/context-engineering.md): Context engineering is the practice of curating the full set of tokens a model sees at each step, including instructions, tools, retrieved data, memory, and conversation history, to maximize the chance of the desired behavior within a limited attention budget. - [Evaluator-Optimizer](https://protologue.com/t/evaluator-optimizer.md): Evaluator-optimizer is a workflow loop in which one model call generates a response and another evaluates it against criteria and provides feedback, repeating until the output passes. - [Meta-Prompting](https://protologue.com/t/meta-prompting.md): Meta-prompting uses a single model as a conductor that breaks a task down and writes prompts for fresh instances of itself acting as specialized experts, then integrates their outputs. - [Orchestrator-Workers](https://protologue.com/t/orchestrator-workers.md): Orchestrator-workers is a pattern in which a central model dynamically breaks a task into subtasks, delegates them to worker model calls or subagents, and synthesizes their results. - [Parallelization](https://protologue.com/t/parallelization.md): Parallelization is a workflow pattern that runs several model calls simultaneously and aggregates their outputs, either by splitting a task into independent sections or by running the same task several times and voting. - [Prompt Chaining](https://protologue.com/t/prompt-chaining.md): Prompt chaining decomposes a task into a fixed sequence of model calls, where each call processes the output of the previous one, often with programmatic checks between steps. - [Routing](https://protologue.com/t/routing.md): Routing is a workflow pattern that classifies an incoming request and directs it to a specialized prompt, tool, or model suited to that category. ## Prompt Optimization - [Automatic Prompt Engineer](https://protologue.com/t/automatic-prompt-engineer.md): Automatic Prompt Engineer (APE) uses a language model to generate candidate instructions for a task from input-output examples, scores each candidate on held-out data, and selects the best one. - [Directional Stimulus Prompting](https://protologue.com/t/directional-stimulus-prompting.md): Directional stimulus prompting trains a small policy model to generate instance-specific hints, such as keywords, that are added to the prompt to steer a large frozen model toward desired outputs. - [DSPy](https://protologue.com/t/dspy.md): DSPy is a framework that treats language-model pipelines as programs of declarative modules, and compiles them by automatically optimizing the prompts and few-shot demonstrations for each module against a metric. - [Emotion Prompting](https://protologue.com/t/emotion-prompting.md): Emotion prompting appends emotional or motivational phrases, such as "This is very important to my career," to a prompt in an attempt to improve model performance. - [Low-Rank Adaptation](https://protologue.com/t/low-rank-adaptation.md): Low-rank adaptation (LoRA) fine-tunes a language model by training small low-rank matrices added to its weight layers while freezing the original weights, drastically reducing the number of trainable parameters. - [Optimization by Prompting](https://protologue.com/t/opro.md): Optimization by PROmpting (OPRO) uses a language model as an optimizer, giving it a meta-prompt containing previously tried prompts and their scores and asking it to propose a better prompt, repeating over many rounds. - [Prefix Tuning](https://protologue.com/t/prefix-tuning.md): Prefix tuning learns continuous task-specific vectors that are prepended to the activations at every layer of a frozen language model, steering generation without changing the model's weights. - [Prompt Caching](https://protologue.com/t/prompt-caching.md): Prompt caching stores the model's processed state for a reused prompt prefix, such as a long system prompt or document, so later requests sharing that prefix are cheaper and faster. - [Prompt Compression](https://protologue.com/t/prompt-compression.md): Prompt compression shortens a prompt by removing tokens that contribute little information, typically scored by a smaller language model, to cut cost and latency while preserving task performance. - [Prompt Tuning](https://protologue.com/t/prompt-tuning.md): Prompt tuning learns a small set of continuous "soft prompt" embeddings that are prepended to the input, by gradient descent, while keeping the language model's weights frozen. ## Reasoning Models & Test-Time Compute - [Extended Thinking](https://protologue.com/t/extended-thinking.md): Extended thinking is a model mode in which the model generates a separate block of reasoning before its final response, with a developer-controlled setting that trades latency and cost for answer quality. - [Reasoning Model](https://protologue.com/t/reasoning-model.md): A reasoning model is a language model trained, typically with reinforcement learning, to produce an extended internal chain of thought before answering, so that it improves with more thinking time on math, coding, and planning tasks. - [Self-Taught Reasoner](https://protologue.com/t/self-taught-reasoner.md): The Self-Taught Reasoner (STaR) bootstraps reasoning ability by having a model generate rationales, keeping those that lead to correct answers, and fine-tuning on them in repeated rounds. - [Test-Time Compute Scaling](https://protologue.com/t/test-time-compute-scaling.md): Test-time compute scaling improves a model's answers by spending more computation at inference, through longer reasoning, more samples, search, or verification, rather than by training a larger model. ## Security & Adversarial Prompting - [Adversarial Suffix](https://protologue.com/t/adversarial-suffix.md): An adversarial suffix is an automatically optimized string of tokens that, when appended to a request, causes an aligned model to comply with requests it would normally refuse. - [Constitutional AI](https://protologue.com/t/constitutional-ai.md): Constitutional AI is a training method in which a model critiques and revises its own outputs according to a written set of principles, and AI-generated preference judgments replace most human labels for harmlessness. - [Indirect Prompt Injection](https://protologue.com/t/indirect-prompt-injection.md): Indirect prompt injection places malicious instructions inside content a model will later retrieve or process, such as a web page, email, or document, so the attack is triggered without the attacker interacting with the model directly. - [Instruction Hierarchy](https://protologue.com/t/instruction-hierarchy.md): The instruction hierarchy is a training approach that teaches a model to prioritize instructions by source, typically system over user over tool output, and to ignore lower-priority instructions that conflict with higher-priority ones. - [Jailbreak](https://protologue.com/t/jailbreak.md): A jailbreak is a prompt crafted to make a model produce outputs its safety training is meant to prevent, often through role-play, hypothetical framing, obfuscation, or other adversarial techniques. - [Many-shot Jailbreaking](https://protologue.com/t/many-shot-jailbreaking.md): Many-shot jailbreaking fills a long context window with many fabricated dialogue examples in which an assistant complies with harmful requests, exploiting in-context learning to override the model's safety training. - [Prompt Injection](https://protologue.com/t/prompt-injection.md): Prompt injection is an attack in which text supplied to a language model, by a user or through data the model processes, contains instructions that override or subvert the instructions of the application's developer. - [Prompt Leaking](https://protologue.com/t/prompt-leaking.md): Prompt leaking is an attack that tricks a model into revealing its hidden system prompt or other confidential instructions. - [Spotlighting](https://protologue.com/t/spotlighting.md): Spotlighting is a family of prompt-level defenses against indirect prompt injection that transform untrusted input, by delimiting, marking every word, or encoding it, so the model can distinguish it from trusted instructions. ## Failure Modes & Evaluation - [Hallucination](https://protologue.com/t/hallucination.md): Hallucination is generated content that is fluent and plausible but unfaithful to the provided source or factually incorrect, such as fabricated citations, facts, or quotations. - [Lost in the Middle](https://protologue.com/t/lost-in-the-middle.md): Lost in the middle is the finding that language models use information at the beginning or end of a long context much more reliably than information placed in the middle. - [Needle in a Haystack](https://protologue.com/t/needle-in-a-haystack.md): Needle in a haystack is a long-context evaluation that hides a specific fact at varying depths in a long distractor document and tests whether the model can retrieve it. - [Prompt Sensitivity](https://protologue.com/t/prompt-sensitivity.md): Prompt sensitivity is the variation in a model's performance caused by superficial changes to a prompt, such as formatting, separators, spacing, or wording, that do not change the task's meaning. - [Sycophancy](https://protologue.com/t/sycophancy.md): Sycophancy is a model's tendency to tailor its answers to match a user's stated beliefs or preferences, including abandoning correct answers when the user pushes back, rather than giving its most accurate response. - [Unfaithful Chain-of-Thought](https://protologue.com/t/unfaithful-chain-of-thought.md): Unfaithful chain-of-thought is stated reasoning that does not reflect the factors that actually determined the model's answer, so the explanation can be plausible yet misleading. ## Optional - [Data downloads](https://protologue.com/data/): JSON, CSV, SKOS JSON-LD and Turtle - [About and citation](https://protologue.com/about/): editorial method, identifiers, license