
이미지: METAL LAB 생성
Summary
- Anthropic let three institutions — Stanford's SALT Lab, Oxford's Human-Centred Computing group, and METR — analyze roughly 250,000 real Claude conversations
- Anthropic's contractual right to intervene was limited to privacy, confidential information, and research accuracy, meaning even unfavorable findings could be published as-is
- Imperial College London ran a separate privacy audit, and the aggregated data was published on Hugging Face
- 파일럿 시기
- 2026년 봄(4~5월 대화 데이터)
- 참여 기관
- 스탠퍼드 SALT 연구소·옥스퍼드 인간정보처리연구소·METR
- 분석 대상
- Claude.ai·Claude Code 대화 약 25만 건
- 분석 도구
- Anthropic Insights(구 Clio)
- 프라이버시 감사
- 임페리얼칼리지런던
- 위반 사례 제외 비율
- 연구별 전체 범주·대화의 5% 미만
- 데이터 공개처
- Hugging Face 데이터셋
- 향후 절차
- 관심 표명 양식 접수 중
Anthropic has given three outside research institutions access to roughly 250,000 real conversations from its AI model, Claude. Stanford University's Social and Language Technologies (SALT) Lab, Oxford University's Human-Centred Computing group, and METR — a nonprofit that evaluates frontier AI models — each independently analyzed this data earlier this spring. Anthropic says its contracts with the three groups explicitly barred the company from influencing what the findings said, letting the researchers publish freely even if the conclusions turned out to be unflattering.
Anthropic Insights, formerly known as Clio
Anthropic has long relied on an internal tool that analyzes millions of anonymized Claude conversations. That tool was originally called "Clio" and has since been renamed "Anthropic Insights." A researcher can pose a question like "What kind of help is this person asking for?" and Claude tags each conversation with an answer, which then gets grouped into categories and reported only as aggregate percentages. No one ever reads the raw conversations — only the final tallies survive — which is by design, as a privacy safeguard.
Anthropic says it had been concerned that data on how people actually use AI was concentrated in the hands of a small number of major AI labs. Outside researchers were left relying on analyses the labs chose to publish themselves, or on public datasets like WildChat that skew heavily toward casual conversation — neither of which was enough to independently verify how AI is really being used. This pilot, granting three institutions access to Anthropic Insights, was meant to close that gap.
What each institution looked at
Between April and May 2026, the three research teams each asked different questions of roughly 250,000 Claude.ai and Claude Code conversations.
Stanford's SALT Lab examined what happens when people and AI work together — what kinds of tasks people hand off to Claude, what role humans keep playing until a task is finished, and where the collaboration tends to break down.
Oxford's Human-Centred Computing group looked at the relationship between how people feel while using Claude and how Claude responds. That study is still being written up, and Anthropic says it will add a link once it's published.
METR is using Claude Code conversations to estimate how much coding agents actually boost productivity in practice, and how that boost changes from one model generation to the next. Anthropic says METR's proposal overlapped with work its own economics research team had already been doing on "agentic coding and the persistence of an expertise premium," so it connected the two teams — a move that ended up helping both projects.
The three research teams at a glance
| Institution | Research question | Status |
|---|---|---|
| Stanford SALT Lab | How human-AI collaboration divides roles, and where it breaks down | Findings published |
| Oxford Human-Centred Computing | Correlation between user emotion and Claude's responses | Report in progress |
| METR | Productivity gains from coding agents, by model generation | Analysis ongoing |

How privacy was protected
Anthropic says it never showed the researchers a single raw conversation. Everything went through the same legal and privacy review as Anthropic's internal research, and only aggregated results were handed over. The contracts also spelled out exactly where Anthropic could intervene: user privacy, information that could help someone violate usage policies, Anthropic's confidential information, and research accuracy. Beyond those four areas — meaning even if a finding was unfavorable to Anthropic — researchers were free to publish it as-is.
The research did surface categories involving violations of Claude's usage policies or terms of service, and Anthropic says it published most of those findings without alteration. The one exception was categories that explained how users had bypassed safety measures, which were filtered out; across the studies, that filtering affected less than 5% of categories and conversations. Separately, Imperial College London conducted an independent privacy audit of the entire data-sharing process to verify that the safeguards actually held up.
What the pilot revealed about its own limits
Anthropic acknowledged that this took far more effort than expected. Its internal research teams can refine a question's wording over several weeks through repeated iteration, but outside partners had to go through a fresh privacy review every time data was shared, making that kind of rapid back-and-forth impractical. As a workaround, researchers were asked to first test their question wording against WildChat, a public dataset — but because WildChat skews toward casual, creative conversations, applying those same categories to real Claude traffic sometimes produced mismatched groupings. Anthropic says it had to write separate interpretation guidelines to help researchers correct for that gap.
By Anthropic's own account, all of this made the pilot slower and more resource-intensive than the pace at which AI labs normally run research internally — a challenge that will need to be solved before the program can expand to more researchers.
How to get involved
Anthropic published the full set of aggregated data from this pilot as a dataset on Hugging Face. It hasn't decided how far to scale the program or how many teams it can support at once going forward; for now, the company says it plans to expand cautiously, prioritizing privacy, safety, and research quality. Anyone interested in conducting research through Anthropic Insights can submit an expression of interest form.
Editor's take
What makes this pilot unusual is that Anthropic deliberately gave up some control over its own narrative. Whenever a company hands data to outside researchers, the first instinct is to assume the results will be cherry-picked for good PR. Anthropic tried to head off that suspicion by writing a clause into its contracts guaranteeing researchers could publish findings even if those findings were unflattering. AI labs have long published their own safety reports, but those reports have always faced a trust problem: the lab itself decided what questions to ask. Handing that question-setting power to outsiders is a real step forward.
Practically speaking, though, this model doesn't look easy to scale. As Anthropic itself admits, the need for a fresh privacy review every time data changes hands means refining even a single question can take weeks, and stand-in validation tools like WildChat don't match real traffic closely enough to avoid introducing errors. Any institution weighing a similar data-sharing effort would do well to standardize two things up front: outsourcing privacy audits to an independent body, and spelling out exactly where the data owner can and can't intervene contractually. Doing that early would save a lot of trial and error later.
The real report card for this experiment will come once Oxford and METR publish their final findings in the coming months — specifically, how much they overlap with, or diverge from, what Anthropic's own internal analysis concluded. A major divergence would be uncomfortable news for Anthropic. But it would also be proof that the program actually worked as intended.




Comments