# LLM-as-a-Judge

> LLM-as-a-judge is the use of a strong language model to grade, score, or compare the outputs of models against criteria, as a scalable substitute for human evaluation.

- Identifier: PTL-0050
- Category: Self-Critique & Verification
- Canonical URL: https://protologue.com/t/llm-as-a-judge/
- Also known as: model-graded evaluation, LLM evaluator, autorater
- Introduced: 2023

## Description

Zheng et al. found strong model judges agreed with human preferences at rates comparable to agreement between humans, while documenting biases toward the first-listed answer, longer answers, and the judge's own outputs.

## Related terms

- [Universal Self-Consistency](https://protologue.com/t/universal-self-consistency/)
- [Evaluator-Optimizer](https://protologue.com/t/evaluator-optimizer/)
- [Process Reward Model](https://protologue.com/t/process-reward-model/)

## Sources

- Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. https://arxiv.org/abs/2306.05685

## Cite this entry

Protologue. (2026). LLM-as-a-Judge. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0050). https://protologue.com/t/llm-as-a-judge/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
