# Test-Time Compute Scaling

> Test-time compute scaling improves a model's answers by spending more computation at inference, through longer reasoning, more samples, search, or verification, rather than by training a larger model.

- Identifier: PTL-0084
- Category: Reasoning Models & Test-Time Compute
- Canonical URL: https://protologue.com/t/test-time-compute-scaling/
- Also known as: inference-time scaling, test-time scaling
- Introduced: 2024

## Description

Snell et al. found that allocating test-time compute adaptively per prompt could be more effective than scaling model parameters for some problems. Self-consistency, best-of-N, tree search, and reasoning models are all forms of test-time scaling.

## Related terms

- [Reasoning Model](https://protologue.com/t/reasoning-model/)
- [Self-Consistency](https://protologue.com/t/self-consistency/)
- [Best-of-N Sampling](https://protologue.com/t/best-of-n-sampling/)
- [Tree of Thoughts](https://protologue.com/t/tree-of-thoughts/)
- [Process Reward Model](https://protologue.com/t/process-reward-model/)

## Sources

- Snell et al. (2024). Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. https://arxiv.org/abs/2408.03314

## Cite this entry

Protologue. (2026). Test-Time Compute Scaling. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0084). https://protologue.com/t/test-time-compute-scaling/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
