protologue

Self-Critique & Verification

Techniques in which a model, or a set of models, checks, critiques, votes on, or revises outputs.

Definitions

Best-of-N Sampling
Best-of-N sampling generates N candidate outputs and returns the one ranked highest by a verifier, reward model, or scoring function.
Chain-of-Verification
Chain-of-Verification (CoVe) reduces hallucination by having the model draft an answer, plan verification questions about its claims, answer those questions independently, and then produce a corrected final answer.
LLM-as-a-Judge
LLM-as-a-judge is the use of a strong language model to grade, score, or compare the outputs of models against criteria, as a scalable substitute for human evaluation.
Mixture-of-Agents
Mixture-of-Agents (MoA) arranges language models in layers, where each model receives all outputs from the previous layer as auxiliary input and an aggregator synthesizes a final response.
Multi-Agent Debate
Multi-agent debate has several model instances propose answers, read each other's reasoning, and revise their answers over multiple rounds until they converge.
Process Reward Model
A process reward model (PRM) scores each intermediate step of a model's reasoning, rather than only the final answer, and is used to select or train better reasoning chains.
Reflexion
Reflexion is an agent technique in which, after a failed attempt, the model writes a verbal reflection on what went wrong and stores it in memory to guide its next attempt.
Self-Refine
Self-Refine is an iterative method in which the same model generates an output, critiques it with specific feedback, and revises it, repeating until a stopping condition is met.