Reinforcement Learning from Human Feedback
Also called RLHF, preference tuning.
Reinforcement learning from human feedback (RLHF) trains a language model to produce outputs people prefer, by learning a reward model from human comparisons and optimizing the model against it.
Description
InstructGPT showed that RLHF made a much smaller model preferred over a far larger base model. RLHF shapes how models respond to prompts, including their helpfulness and refusals, and is linked to failure modes such as sycophancy. Direct preference optimization is a widely used simpler alternative.
Sources
- Ouyang et al. (2022). Training language models to follow instructions with human feedback.
Cite this entry
Protologue. (2026). Reinforcement Learning from Human Feedback. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0011). https://protologue.com/t/rlhf/
BibTeX
@misc{protologue_rlhf,
title = {Reinforcement Learning from Human Feedback},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0011},
url = {https://protologue.com/t/rlhf/}
}