# Direct Preference Optimization

> Direct preference optimization (DPO) aligns a language model to human preferences by training directly on preferred-versus-rejected response pairs, without fitting a separate reward model or running reinforcement learning.

- Identifier: PTL-0012
- Category: Foundations
- Canonical URL: https://protologue.com/t/direct-preference-optimization/
- Also known as: DPO
- Introduced: 2023

## Description

DPO reframes the RLHF objective as a simple classification-style loss over preference pairs. It became a common alternative to RLHF because it is stable and cheap to run.

## Broader terms

- [Reinforcement Learning from Human Feedback](https://protologue.com/t/rlhf/)

## Sources

- Rafailov et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. https://arxiv.org/abs/2305.18290

## Cite this entry

Protologue. (2026). Direct Preference Optimization. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0012). https://protologue.com/t/direct-preference-optimization/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
