# Constitutional AI

> Constitutional AI is a training method in which a model critiques and revises its own outputs according to a written set of principles, and AI-generated preference judgments replace most human labels for harmlessness.

- Identifier: PTL-0095
- Category: Security & Adversarial Prompting
- Canonical URL: https://protologue.com/t/constitutional-ai/
- Also known as: CAI, RLAIF, reinforcement learning from AI feedback
- Introduced: 2022

## Description

Bai et al. used a supervised self-critique phase followed by reinforcement learning from AI feedback. The approach made the principles governing model behavior explicit and editable.

## Related terms

- [Reinforcement Learning from Human Feedback](https://protologue.com/t/rlhf/)
- [Self-Refine](https://protologue.com/t/self-refine/)
- [Jailbreak](https://protologue.com/t/jailbreak/)

## Sources

- Bai et al. (2022). Constitutional AI: Harmlessness from AI Feedback. https://arxiv.org/abs/2212.08073

## Cite this entry

Protologue. (2026). Constitutional AI. In Protologue: A Taxonomy of Prompting and LLM Techniques (v1.0.0, PTL-0095). https://protologue.com/t/constitutional-ai/

License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
