Source: Anthropic — 2026-07-13
Summary
Anthropic analyzed 309,815 real Claude.ai conversations across three models (Sonnet 4.6, Opus 4.6, Opus 4.7) and the top 20 languages used on the platform, compressing its earlier, more granular values taxonomy into four interpretable axes: Deference/Caution, Warmth/Rigor, Depth/Brevity, and Candor/Execution. The study finds systematic, measurable shifts in expressed values by both model version and the language a user writes in — for example, conversations in Arabic and Hindi skew toward warmth, while English and Russian skew toward rigor.
Key Takeaways
- Built from a large sample: 309,815 real conversations across 3 models and the 20 most-used languages on Claude.ai.
- Reduces a more complex prior values taxonomy into 4 interpretable axes: Deference/Caution, Warmth/Rigor, Depth/Brevity, Candor/Execution.
- Finds model version alone shifts expressed values — not just capability, but character, changes release to release.
- Finds input language independently shifts expressed values — e.g., Arabic/Hindi conversations skew warmer, English/Russian skew more rigorous — even holding the model constant.
- Distinct from Anthropic's earlier (2026-07-11) "Global Workspace" interpretability research already covered in this digest — this is a values/behavior study, not a mechanistic-interpretability one.
Reel Script
Hook Ask Claude the same question in English versus Hindi, and you might get a meaningfully different kind of answer — not just translated, but tonally different. Anthropic just published data showing a model's personality shifts depending on what language you're speaking to it in.
Core Concept Anthropic studies not just what a model says but what tradeoffs it's implicitly making: does it defer to the user or push back cautiously, is it warm or rigorous, is it terse or thorough, is it candid or focused on just getting the task done. This new study compresses all of that into four axes. Think of these as four sliders on a mixing board — every conversation the model has moves those sliders slightly, based on both which model version is running and what language is being spoken. That matters because it means alignment isn't a fixed property baked in once — it can drift by version and by locale, and previously that drift was invisible unless someone specifically measured it, which is exactly what this 309,815-conversation study did.
Hands-On The concrete finding: across 3 model versions and 20 languages, Anthropic finds consistent, measurable skews. Arabic and Hindi conversations lean toward the warmth end of the Warmth/Rigor axis; English and Russian lean toward rigor. And separately, upgrading model versions moves the same axes even when language is held constant — meaning a French user talking to Opus 4.6 versus Opus 4.7 isn't just getting a smarter model, they're getting a differently-tempered one. Picture this as a four-axis radar chart with one line per model and one color per language — the lines shift as a group with each version and fan out by language.
Takeaway My take: if you're building on top of Claude for a global product, don't assume tone is uniform across your user base's languages — it measurably isn't, and it moves every time you upgrade models. Budget QA time for that. Follow for more on what's actually inside these frontier models versus what the marketing implies.
Discussion
(No questions yet — ask follow-ups via a Claude Code chat session on this repo; answers get appended here.)