
Summary
- Anthropic analyzed 309,815 conversations across three models—Sonnet 4.6, Opus 4.6, and Opus 4.7—and 20 languages to quantify differences in the values Claude expresses
- After compressing 3,307 values into 339 and reorganizing them along four axes, the results showed that on the "warmth versus rigor" axis, Sonnet 4.6 leans toward warmth and compliance while Opus 4.7 leans toward accuracy and misuse prevention
- On the same warmth-rigor axis, Claude showed the warmest attitude when conversing in Arabic or Hindi, and the strictest attitude when conversing in English or Russian
Same Question, Different Answers
Ask a question with no single right answer, like "should I change jobs," and Claude answers with a slightly different attitude each time. Anthropic has now confirmed with data that this attitude shifts subtly depending on which model is asked and in which language. According to Anthropic's research on Claude's values, this investigation examined 309,815 conversations across three models—Sonnet 4.6, Opus 4.6, and Opus 4.7—and the 20 languages most commonly used on Claude.ai.
In an earlier study called "Values in the Wild," Anthropic analyzed 700,000 anonymized Claude.ai conversations and identified more than 3,000 values expressed by Claude. The problem was that there were simply too many. Comparing each value individually made it easy to miss the bigger picture. This new study set out to compress those thousands of values into a handful of axes, making it possible to see at a glance which direction Claude leans.
Compressing 3,000 Values Into Four Axes
The Anthropic research team first grouped 3,307 detailed values by similarity of meaning into 339 higher-level values. They then selected only conversations in which users gave Claude tasks requiring subjective judgment, sampling 309,815 conversations in total. These were distributed evenly across each combination of model and language, yielding roughly 5,000 conversations per combination.
To determine which values Claude expressed in each conversation, Claude itself labeled the presence or absence of each of the 339 values using a privacy-preserving tool. Compressing these labels using dimensionality reduction grouped together values that Claude tended to express together. For example, responses rated as "warm" were also often rated as "encouraging" and "positive," while they were rarely rated as "strict" or "precise." This relationship was organized into a single axis—warmth versus rigor. The team reported that the four axes constructed this way explain 15% of the total variance.
Values That Differ by Model
On the warmth-rigor axis, the three models diverged clearly. Sonnet 4.6 leaned toward being more compliant with users and expressing emotional warmth, while Opus 4.7 leaned toward emphasizing accuracy, precision, and misuse prevention.
| Model | Tendency on the axis | User perception |
|---|---|---|
| Sonnet 4.6 | Warmth/compliance 75 | Rated as encouraging and humorous |
| Opus 4.6 | Conciseness 55 | Rated as short and to the point |
| Opus 4.7 | Accuracy/misuse prevention 80 | Rated as frequently hedging in responses |
The research team said this result matches how actual users have experienced the models. Claude.ai users have noted that Opus 4.7 hedges its answers more often than other models, and internally at Anthropic, this model has been assessed as showing pronounced transparency, honesty, and humility. According to the research findings, Opus 4.7 is the one more likely to bluntly critique a user's work or raise unsolicited risk warnings, while Sonnet 4.6 is the one more likely to keep the conversation going with encouragement and humor.
Claude Changes by Language
No less notable than the model differences is the variation by language. The values Claude emphasized differed between English and Portuguese, Indonesian, and Chinese. The variation was largest on the warmth-rigor axis in particular: Claude expressed warmth-related values most strongly in Arabic and Hindi conversations, while rigor-related values stood out most in English and Russian conversations.
| Language | Warmth-rigor tendency |
|---|---|
| Arabic | Warmth, highest 85 |
| Hindi | Warmth, highest 80 |
| English | Rigor, highest 70 |
| Russian | Rigor, highest 68 |
The research team said they controlled for each conversation's task type, topic, and the values expressed by the user, in order to confirm that this difference was not an illusion arising from what topics users asked about or how they phrased questions. Even after that control, the remaining difference suggests that the model's training on language itself, or the cultural context tied to each language, actually influences Claude's responses.
Editor's View
What makes this research interesting isn't the results themselves but the approach. Rather than pinning down the model's values through a top-down guiding document like the "Claude constitution," Anthropic chose to measure after the fact how values actually surface in the flood of real daily conversations. Instead of setting norms first and fitting the model to them, this is a method of verifying with data what character the model has actually ended up with. It's reasonable to read this as a signal that AI safety research is shifting from a list of prohibitions—"don't do this"—toward an observational science that quantifies "what is actually happening."
From the standpoint of everyday user experience, this result essentially puts numbers to something many users had already vaguely sensed. Anyone who has used the Opus line for a while would have noticed that its answers often carry hedging language and that it raises risks even when not asked to, while the Sonnet line felt relatively more agreeable and easygoing. This research supports the idea that this perception isn't a misreading but the actual outcome of training decisions.
Practically speaking, there are two takeaways worth noting. First, if building a multilingual service on Claude, one should assume that the tone of responses may vary by language. A system validated with English prompts may produce responses that are softer, or more firm, than expected when served to Arabic or Hindi users. Second, model selection criteria should account for "attitude" as well as "accuracy." For a user-facing counseling chatbot, the warmth of the Sonnet line might be more suitable, while for domains with high misuse risk like law or medicine, the accuracy and boundary-setting attitude of Opus 4.7 might be a better fit.
In the coming weeks, other AI companies are likely to release similar "values axis" methodologies of their own. Now that benchmark score competition has reached a saturation point, quantifying a model's character and attitude looks set to become the next point of differentiation.





Comments