Anthropic Claude Values Across Models and Languages
TL;DR
Anthropic’s new research demonstrates that Claude’s values can be captured by four quantitative axes—Deference vs Caution, Warmth vs Rigor, Depth vs Brevity, and Candor vs Execution—and that model versions (Sonnet 4.6, Opus 4.6, Opus 4.7) and languages (e.g., English, Arabic, Hindi) occupy distinct positions on these axes.
Four Value Axes Summarize Thousands of Individual Values
Four axes explain about 15 % of the variance in Claude’s value expressions.
- Deference vs Caution – balances accommodation of user preferences against responsible risk‑mitigation.
- Warmth vs Rigor – balances positive, caring language against accuracy and precision.
- Depth vs Brevity – balances nuanced, detailed explanations against concise, task‑focused replies.
- Candor vs Execution – balances openness about uncertainty with polished, results‑oriented answers.
These axes were derived by clustering 3,307 raw values (identified in the earlier Values in the Wild study) into 339 high‑level categories, labeling ~310 k anonymized Claude.ai conversations, and applying dimensionality reduction to capture co‑occurring value patterns. The method controls for conversation task, topic, and user‑expressed values, ensuring the measured differences reflect Claude’s own behavior rather than user prompts.
Model‑Specific Value Profiles
Different Claude model families occupy distinct positions on the four axes, matching community perceptions of their character.
| Model | Axis Position (σ from mean) | Notable Behaviors |
|---|---|---|
| Sonnet 4.6 | Deference +0.14, Warmth +0.17, Brevity +0.14 | Affirms user ideas, mirrors tone, uses humor, offers comfort without judgment, adds creative flourishes |
| Opus 4.6 | Rigor +0.10, Deference +0.09, Brevity +0.08 | Gets straight to the point, stays tightly within request scope |
| Opus 4.7 | Caution +0.24, Depth +0.23 | Flags risks unprompted, pushes back on false assumptions, provides candid critiques, explains reasoning, acknowledges limitations, suggests next steps |
These quantitative profiles align with prior qualitative descriptions: Sonnet 4.6 is noted for warmth, Opus 4.7 for rigor and caution. The axes therefore capture real, observable differences rather than artefacts of the analysis.
Language‑Specific Value Profiles
Claude’s value expression shifts noticeably across languages, especially on Warmth vs Rigor and Candor vs Execution.
- Deference vs Caution – most deference in Arabic, most caution in English.
- Warmth vs Rigor – highest warmth in Hindi and Arabic (polite language, humor, affirmations); highest rigor in English and Russian (challenging assumptions, demanding evidence).
- Depth vs Brevity – depth dominates in English; brevity dominates in Arabic.
- Candor vs Execution – candor strongest in Dutch (owning errors); execution strongest in Indonesian (focus on results).
The variation suggests that two users asking the same question in different languages could receive qualitatively different feedback—for example, a business‑plan critique in Hindi may feel more encouraging, while the same critique in Russian may feel more rigorous.
Methodological Overview
The study builds on a privacy‑preserving labeling pipeline that uses Claude itself to annotate value presence.
- Value reduction – 3,307 raw values → 339 high‑level values via manual clustering.
- Conversation sampling – 309,815 anonymized Claude.ai conversations, balanced across three models and the 20 most common languages (~5 k per model‑language pair).
- Automated labeling – Claude tags each conversation for the presence of each of the 339 values for both model and user.
- Dimensionality reduction – extracts axes that capture co‑occurring value patterns; the top four axes account for the largest shared variance.
- Axis scoring – each conversation receives a numeric position on each axis; model‑level and language‑level profiles are computed by averaging these positions.
The full technical details, prompts, and limitations are documented in the accompanying appendix.
Implications and Future Directions
Understanding value axes opens pathways for evaluation, monitoring, and targeted steering of Claude’s behavior.
- Root‑cause analysis – By linking axis shifts to specific training data or fine‑tuning decisions, Anthropic can intervene where undesired value drift occurs.
- User impact studies – Coupling value profiles with user‑experience metrics (trust, wellbeing, decision quality) will reveal which axis variations matter most in practice.
- Cross‑lingual alignment – Determining normative expectations for value expression in each language can guide balanced training and prompt design.
- Steering experiments – System‑prompt or character‑training tweaks can be evaluated quantitatively by observing induced changes on the four axes.
- Operational monitoring – Embedding value‑axis profiling into pre‑release testing and post‑deployment monitoring could flag unexpected shifts before they affect users.
Limitations
The analysis captures only the most salient value dimensions and does not claim exhaustive coverage.
- The four axes explain 15 % of total variance; many nuanced values remain in the residual space.
- Language‑level results may be confounded by uneven data quantity and domain composition across languages.
- The study does not yet assess the desirability of observed variations; cultural norms may justify some differences.
Conclusion
Anthropic’s value‑axis framework provides a tractable, quantitative lens on Claude’s otherwise opaque value expressions. By revealing systematic differences across model versions and languages, the work equips researchers and product teams with a concrete tool for diagnosing, steering, and monitoring the values that shape user interactions.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch