METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Claude Answers Strictly as Opus, Warmly as Sonnet

Anthropic analyzed over 300,000 conversations and found that Claude expresses different values depending on the model and language

Claude Answers Strictly as Opus, Warmly as Sonnet

Summary

  • Anthropic analyzed 309,815 conversations across three models—Sonnet 4.6, Opus 4.6, and Opus 4.7—and 20 languages to quantify differences in the values Claude expresses
  • After compressing 3,307 values into 339 and reorganizing them along four axes, the results showed that on the "warmth versus rigor" axis, Sonnet 4.6 leans toward warmth and compliance while Opus 4.7 leans toward accuracy and misuse prevention
  • On the same warmth-rigor axis, Claude showed the warmest attitude when conversing in Arabic or Hindi, and the strictest attitude when conversing in English or Russian

Same Question, Different Answers

Ask a question with no single right answer, like "should I change jobs," and Claude answers with a slightly different attitude each time. Anthropic has now confirmed with data that this attitude shifts subtly depending on which model is asked and in which language. According to Anthropic's research on Claude's values, this investigation examined 309,815 conversations across three models—Sonnet 4.6, Opus 4.6, and Opus 4.7—and the 20 languages most commonly used on Claude.ai.

In an earlier study called "Values in the Wild," Anthropic analyzed 700,000 anonymized Claude.ai conversations and identified more than 3,000 values expressed by Claude. The problem was that there were simply too many. Comparing each value individually made it easy to miss the bigger picture. This new study set out to compress those thousands of values into a handful of axes, making it possible to see at a glance which direction Claude leans.

Compressing 3,000 Values Into Four Axes

The Anthropic research team first grouped 3,307 detailed values by similarity of meaning into 339 higher-level values. They then selected only conversations in which users gave Claude tasks requiring subjective judgment, sampling 309,815 conversations in total. These were distributed evenly across each combination of model and language, yielding roughly 5,000 conversations per combination.

To determine which values Claude expressed in each conversation, Claude itself labeled the presence or absence of each of the 339 values using a privacy-preserving tool. Compressing these labels using dimensionality reduction grouped together values that Claude tended to express together. For example, responses rated as "warm" were also often rated as "encouraging" and "positive," while they were rarely rated as "strict" or "precise." This relationship was organized into a single axis—warmth versus rigor. The team reported that the four axes constructed this way explain 15% of the total variance.

Values That Differ by Model

On the warmth-rigor axis, the three models diverged clearly. Sonnet 4.6 leaned toward being more compliant with users and expressing emotional warmth, while Opus 4.7 leaned toward emphasizing accuracy, precision, and misuse prevention.

ModelTendency on the axisUser perception
Sonnet 4.6Warmth/compliance 75Rated as encouraging and humorous
Opus 4.6Conciseness 55Rated as short and to the point
Opus 4.7Accuracy/misuse prevention 80Rated as frequently hedging in responses

The research team said this result matches how actual users have experienced the models. Claude.ai users have noted that Opus 4.7 hedges its answers more often than other models, and internally at Anthropic, this model has been assessed as showing pronounced transparency, honesty, and humility. According to the research findings, Opus 4.7 is the one more likely to bluntly critique a user's work or raise unsolicited risk warnings, while Sonnet 4.6 is the one more likely to keep the conversation going with encouragement and humor.

Claude Changes by Language

No less notable than the model differences is the variation by language. The values Claude emphasized differed between English and Portuguese, Indonesian, and Chinese. The variation was largest on the warmth-rigor axis in particular: Claude expressed warmth-related values most strongly in Arabic and Hindi conversations, while rigor-related values stood out most in English and Russian conversations.

LanguageWarmth-rigor tendency
ArabicWarmth, highest 85
HindiWarmth, highest 80
EnglishRigor, highest 70
RussianRigor, highest 68

The research team said they controlled for each conversation's task type, topic, and the values expressed by the user, in order to confirm that this difference was not an illusion arising from what topics users asked about or how they phrased questions. Even after that control, the remaining difference suggests that the model's training on language itself, or the cultural context tied to each language, actually influences Claude's responses.

Editor's View

What makes this research interesting isn't the results themselves but the approach. Rather than pinning down the model's values through a top-down guiding document like the "Claude constitution," Anthropic chose to measure after the fact how values actually surface in the flood of real daily conversations. Instead of setting norms first and fitting the model to them, this is a method of verifying with data what character the model has actually ended up with. It's reasonable to read this as a signal that AI safety research is shifting from a list of prohibitions—"don't do this"—toward an observational science that quantifies "what is actually happening."

From the standpoint of everyday user experience, this result essentially puts numbers to something many users had already vaguely sensed. Anyone who has used the Opus line for a while would have noticed that its answers often carry hedging language and that it raises risks even when not asked to, while the Sonnet line felt relatively more agreeable and easygoing. This research supports the idea that this perception isn't a misreading but the actual outcome of training decisions.

Practically speaking, there are two takeaways worth noting. First, if building a multilingual service on Claude, one should assume that the tone of responses may vary by language. A system validated with English prompts may produce responses that are softer, or more firm, than expected when served to Arabic or Hindi users. Second, model selection criteria should account for "attitude" as well as "accuracy." For a user-facing counseling chatbot, the warmth of the Sonnet line might be more suitable, while for domains with high misuse risk like law or medicine, the accuracy and boundary-setting attitude of Opus 4.7 might be a better fit.

In the coming weeks, other AI companies are likely to release similar "values axis" methodologies of their own. Now that benchmark score competition has reached a saturation point, quantifying a model's character and attitude looks set to become the next point of differentiation.

Comments