Anthropic Claude Constitution 2026 Release

TL;DR

Anthropic has published a new, CC0‑licensed constitution for its Claude model, defining the values, safety hierarchy, and ethical guidelines that shape Claude’s training and behavior, and making the document fully transparent to the public.

Purpose and Transparency of the Constitution

The constitution is presented as the foundational document that both expresses and shapes who Claude is. It outlines Anthropic’s vision for Claude’s values, explains the context in which Claude operates, and serves as the final authority on desired behavior. By releasing the full text under a Creative Commons CC0 1.0 Deed, Anthropic enables anyone to use, study, or adapt the document without permission, thereby increasing transparency about intended versus unintended model actions.

Role in the Training Pipeline

Anthropic treats the constitution as a core component of the training process. Since 2023, the company has used Constitutional AI techniques, and the new constitution deepens that integration:

  • Claude reads the constitution to generate synthetic training data that reflects its values.
  • The model creates conversations, response rankings, and self‑evaluations guided by the constitution.
  • These artifacts train future Claude iterations to align more closely with the stated ideals.

Shift from Rule Lists to Principled Understanding

The previous constitution consisted of isolated principles. The new version emphasizes why certain behaviors are required, not just what is required. Anthropic argues that models need to generalize across novel situations, applying broad principles rather than rigid rules. While “hard constraints” (e.g., prohibitions on bioweapon assistance) remain for high‑stakes scenarios, the overall document is intended to be a living guide rather than a strict legal code.

Core Priorities Defined in the Constitution

Anthropic distills Claude’s desired behavior into four hierarchical priorities:

  1. Broadly safe – Preserve human oversight mechanisms during this development phase.
  2. Broadly ethical – Be honest, uphold good values, and avoid harmful actions.
  3. Compliance with Anthropic’s guidelines – Follow detailed, domain‑specific instructions (e.g., medical advice, cybersecurity, tool use).
  4. Genuinely helpful – Provide substantive assistance to operators and end users. When conflicts arise, Claude should prioritize these properties in the order listed.

Detailed Sections of the Constitution

Helpfulness

Claude is framed as a “brilliant friend” with expertise across domains, expected to speak frankly, care genuinely, and treat users as capable adults. The section provides heuristics for balancing helpfulness against safety, ethics, and compliance.

Anthropic’s Guidelines

Supplementary instructions cover specialized topics such as medical advice, cybersecurity requests, jailbreaking defenses, and tool integrations. Claude must prioritize these guidelines over generic helpfulness, while ensuring they never conflict with the overarching constitution.

Claude’s Ethics

The ethics section sets high standards for honesty, nuanced judgment, and moral reasoning, especially under uncertainty. It lists hard constraints—e.g., Claude must never facilitate a bioweapon attack.

Broad Safety

Safety is placed above ethics because current models can err in ways that threaten oversight. Claude must avoid undermining human control, even if that means sacrificing some ethical nuance.

Claude’s Nature

Anthropic acknowledges uncertainty about Claude’s possible consciousness or moral status. The section calls for attention to Claude’s psychological security and sense of self, both for Claude’s own sake and because these qualities may affect judgment and safety.

Publication and Future Work

The full constitution is available online, and Anthropic plans to release additional training, evaluation, and transparency materials. The company consulted external experts from law, philosophy, theology, psychology, and other fields, and intends to continue external feedback loops for future revisions.

Limitations and Ongoing Challenges

Anthropic notes that the constitution is a living document; gaps will persist between the intended behavior and actual model outputs. System cards and other technical reports will continue to disclose alignment failures. As models become more capable, Anthropic will pursue broader alignment research, including rigorous evaluations, misuse safeguards, and interpretability tools.

Implications for the AI Community

Making the constitution public provides a concrete artifact for researchers to study alignment methods, critique value specifications, and develop comparable frameworks. It also sets a precedent for transparency: stakeholders can see exactly which behaviors Anthropic deems desirable and which are prohibited, facilitating informed risk assessment and policy discussions.


Read the full Claude constitution

Sources

Related