Process Reward Model
Also called PRM, step-level verifier, process supervision.
A process reward model (PRM) scores each intermediate step of a model's reasoning, rather than only the final answer, and is used to select or train better reasoning chains.
Description
Lightman et al. showed process supervision outperformed outcome supervision for selecting correct solutions to competition math problems, and released a large dataset of step-level human labels.
Sources
- Lightman et al. (2023). Let's Verify Step by Step.
Cite this entry
Protologue. (2026). Process Reward Model. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0053). https://protologue.com/t/process-reward-model/
BibTeX
@misc{protologue_process_reward_model,
title = {Process Reward Model},
author = {{Protologue}},
year = {2026},
howpublished = {Protologue: A Taxonomy of Prompting and LLM Techniques, v1.0.0},
note = {Entry PTL-0053},
url = {https://protologue.com/t/process-reward-model/}
}