Claude's Constitution: Implementing Constitutional AI for Model Alignment
Anthropic has developed "Constitutional AI" (CAI), a training approach that gives language models explicit values via a written constitution. This method allows the AI to determine appropriate engagement and behavioral boundaries based on transparent principles, making the system's values easier to understand and adjust than those derived implicitly from large-scale human feedback.
The Limitations of Human-Centric Feedback
Traditional model alignment often relies on human contractors comparing model outputs to select the "better" response based on principles like helpfulness or harmlessness. Anthropic identifies three primary shortcomings to this approach:
- Human Exposure: The process may require human reviewers to interact with disturbing or traumatic content.
- Scalability: As models produce more complex responses, it becomes difficult for human crowdworkers to keep up with or fully comprehend the outputs.
- Resource Intensity: The time and resources required for reviewing subsets of outputs make the process inaccessible to many researchers.
How Constitutional AI Works
Constitutional AI replaces human supervision with AI feedback to evaluate outputs. The system is guided by a constitution—a set of normative principles—to avoid toxic or discriminatory outputs and prevent assistance in illegal or unethical activities.
The CAI training process occurs in two distinct phases:
- Critique and Revision: The model is trained to critique and revise its own responses using the constitutional principles and a few provided examples.
- Reinforcement Learning (RL): Instead of using human feedback, the model is trained via RL using AI-generated feedback based on the constitution to select the more harmless output.
Anthropic reports that CAI can produce a "Pareto improvement," where the model becomes both more helpful and more harmless than models trained via reinforcement learning from human feedback (RLHF). In testing, CAI-trained models responded more appropriately to adversarial inputs without becoming evasive, achieving these harmlessness results purely through AI supervision.
Composition of Claude's Constitution
Claude's constitution is an evolving set of principles drawn from diverse sources to ensure a broad and inclusive value system. These sources include:
- The UN Declaration of Human Rights: Used to cover broad, core human values.
- Global Platform Guidelines: Inspired by Apple's Terms of Service to address modern digital issues like data privacy and online impersonation.
- Frontier AI Research: Incorporating best practices from other labs, such as DeepMind's Sparrow Principles.
- Non-Western Perspectives: Specific principles designed to ensure the model is not solely reflective of Western, rich, or industrialized cultures.
- Empirical Research: Principles developed through trial-and-error to optimize generalization and effectiveness.
Balancing Ethics and User Experience
Anthropic discovered that CAI-trained models could occasionally become "judgmental or annoying." To counter this, they integrated principles that encourage proportionate responses, instructing the model to be ethical and moral without sounding "excessively condescending, reactive, obnoxious, or condemnatory."
Implementation and Future Directions
During training, the model does not reference every principle for every response. Instead, it pulls one principle at random during the supervised learning phase (critique and revision) and the reinforcement learning phase (evaluation).
Anthropic views Constitutional AI as a way to make AI value systems explicit and alterable. The lab is exploring more democratic methods for producing the constitution and the possibility of offering customizable constitutions for specific use cases to avoid the imposition of a single political ideology or viewpoint.
Sources
- OriginalClaude’s Constitution
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch