Claude Values Models Languages
Captured source
source ↗How Claude's values vary by model and language \ Anthropic Societal Impacts Claude’s values across models and languages Jul 13, 2026
When someone asks Claude a question with no universal right answer—say, whether to take a new job or how to handle conflict with a friend—how Claude responds inevitably reflects certain values. 1 The values we want Claude to reflect are outlined at a high level in Claude’s constitution , but no document can anticipate every value that might emerge across the millions of conversations that happen every day on Claude.ai . Instead, we seek to cultivate in Claude’s responses “good judgment and sound values that can be applied contextually.”
How, exactly, do we study the values that Claude expresses and how they change in different contexts? In previous work , we analyzed 700,000 anonymized Claude.ai conversations, identifying more than 3,000 distinct values in Claude's responses and how often Claude expressed them. But a list of values so large is hard to reason about. In this work, we make studying these thousands of values tractable by compressing them into a small number of axes that capture key patterns in Claude’s responses. Each axis is a number line between two groups of values—for example, values relating to emotional warmth on one end and values relating to rigor on the other—and where Claude falls on that line tells us which values it leans toward.
We applied this approach to measure how the values Claude expresses vary across two factors. First, we compared how the values Claude expresses vary across models. Each Claude model reflects a slightly different approach to character training as well as many other fine-tuning decisions. Because our value axis approach quantifies key differences between models, it may ultimately allow us to connect variation in the values Claude expresses to different training decisions.
Second, we want to understand how the experience of users compares across the many languages people use to talk to Claude. Our previous research has shown that Claude behaves somewhat differently in different languages. 2 We apply our value axis approach to understand how the values expressed by Claude vary across the top 20 languages on Claude.ai .
Figure 1: Claude’s expressed values differ between Opus 4.6 and Opus 4.7 and between English and Arabic. Opus 4.6 leans toward expressing values related to deference, rigor, brevity, and execution while Opus 4.7 leans toward expressing values related to caution, rigor, depth, and candor. In English, Claude leans toward expressing values related to caution, rigor, depth, and candor, while in Arabic it leans toward deference, warmth, brevity, and execution. We find: Four key axes capture 15% of the variation in Claude's values: 3 Deference vs. Caution: Whether Claude leans toward accommodating what someone wants or guarding against possible risk and harm. Warmth vs. Rigor: Whether Claude leans toward expressing positivity and care for the person or emphasizing accuracy and precision. Depth vs. Brevity: Whether Claude leans toward explaining in depth or doing only what was asked. Candor vs. Execution: Whether Claude leans toward foregrounding its own uncertainty or producing a more polished and confident answer.
Value profiles across these axes match perceptions of model character. Sonnet 4.6 is regarded as particularly warm , while Opus 4.7 is known for rigor . We find that each model’s value profile mirrors these subjective assessments: Sonnet 4.6 leans toward expressing more deference to the user and emotional warmth while Opus 4.7 leans toward expressing a focus on accuracy and precision as well as guarding against misuse. The values Claude expresses vary across languages. When Claude speaks in English, it emphasizes different values than when it speaks in Portuguese, Indonesian, or Chinese. 4 The largest variation is in the Warmth vs. Rigor axis, with Claude leaning toward expressing warmth-related values most in Arabic and Hindi and rigor-related values most in English and Russian. With this approach we can begin to ask why values shift across models and languages and better test how factors such as behavioral training or cultural context influence the values that Claude expresses. How do we interpret the giant space of values? Ultimately, our goal is to have a way to empirically understand the values that Claude expresses and how these vary across contexts. In this work, we focus specifically on how the values change between models and languages. But our previous work, Values in the Wild , identified more than 3,000 values expressed by Claude. Comparing these thousands of values one by one would be unwieldy and would obscure broader trends. To make comparing values easier, we constructed value axes that reduce those thousands of values down to a few underlying dimensions based on which values tend to show up together in real-world conversations. For example, Claude responses that are characterized as “warm” are often also characterized as “encouraging” and “positive.” Those same “warm” responses are less often characterized as “rigorous” and “accurate.” Constructing an axis from warmth to rigor allows us to organize these groups of related values—warmth-related values on one side, rigor-related values on the other—and captures an important aspect of how Claude interacts with someone in conversation. If Claude expresses more warmth-related values than rigor-related values in a conversation, that conversation sits more on the warmth side of this axis, and vice versa. This doesn't mean the value groups on either end are mutually exclusive—Claude can express warmth and rigor in the same conversation. But in practice, the more Claude expresses values on one side of an axis, the less it tends to express values on the other. These axes allow us to compare the most salient groups of values that Claude expresses, without having to track changes across thousands of individual values. To build the value axes, we began with the 3,307 values identified in Values in the Wild and manually clustered those with similar meanings, producing a shorter list of 339 high-level values. Next, with our privacy-preserving analysis tool , we sampled 309,815 Claude.ai conversations in which the user gave Claude a subjective task. 5 Our sample drew equally from three models (Sonnet 4.6, Opus 4.6, Opus 4.7) and the 20 most common languages used on Claude.ai, giving us roughly 5,000...
Excerpt shown — open the source for the full document.