Constitutional AI is Anthropic's alignment method in which a model is trained against an explicit set of principles (a 'constitution'): the model critiques and
Anthropic's alignment method in which a model is trained against an explicit set of principles (a 'constitution'): the model critiques and revises its own outputs according to those principles, and AI feedback replaces much of the human labeling in preference training.
Its significance is scalability and transparency: alignment is guided by inspectable written principles rather than thousands of implicit human judgments. It shaped how the industry thinks about steering model values.