protologue

Reinforcement Learning from Human Feedback

Also called RLHF, preference tuning.

Reinforcement learning from human feedback (RLHF) trains a language model to produce outputs people prefer, by learning a reward model from human comparisons and optimizing the model against it.

Description

InstructGPT showed that RLHF made a much smaller model preferred over a far larger base model. RLHF shapes how models respond to prompts, including their helpfulness and refusals, and is linked to failure modes such as sycophancy. Direct preference optimization is a widely used simpler alternative.

Sources

  1. Ouyang et al. (2022). Training language models to follow instructions with human feedback.

Cite this entry

Protologue. (2026). Reinforcement Learning from Human Feedback. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0011). https://protologue.com/t/rlhf/

BibTeX
@misc{protologue_rlhf,
  title = {Reinforcement Learning from Human Feedback},
  author = {{Protologue}},
  year = {2026},
  howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
  note = {Entry PTL-0011},
  url = {https://protologue.com/t/rlhf/}
}

Markdown JSON