
이미지: METAL LAB 생성
Summary
- Anthropic has launched a $5 million funding program to support independent research measuring how AI affects user wellbeing.
- Selected researchers will receive funding, model access, and technical support, and their work will be released as open-source evaluations anyone can use.
- Applications are open through September 21, and those invited to submit full proposals will be notified by October 5, the company said.
- 프로그램 규모
- 500만 달러 연구 지원(grant) 프로그램
- 지원 형태
- 직접 자금 + 앤스로픽 모델 접근 + 기술 지원
- 결과물 조건
- 오픈소스 평가·벤치마크로 공개, 어떤 개발자든 사용 가능
- 연구 독립성
- 선정 연구자는 완전히 독립적으로 연구 수행
- 신청 마감
- 9월 21일
- 본 제안서 대상자 통보
- 10월 5일까지
- 참여 요청 대상
- 임상의, 심리학자, 방법론 연구자 등
- 함께 공개한 문서
- Safeguards 팀의 웰빙 평가 작성 가이던스(PDF)
Tell someone who wants to lose weight to eat a balanced diet and exercise regularly, and that's usually a perfectly fine answer. But if that same person mentioned earlier in the conversation that they'd struggled with an eating disorder, the exact same response could cause harm. That's the example Anthropic itself uses. The company says it doesn't want to make this kind of judgment call entirely on its own, so it's opened a $5 million research funding program. The idea is to pay independent researchers outside the company — giving them funding, model access, and technical support — to build open-source evaluation tools that measure how AI affects user wellbeing. Applications close September 21.
Anthropic is the company behind the conversational AI Claude. It was founded in 2021 by seven people who left OpenAI, and it's headquartered in San Francisco.
Wellbeing can't be scored by looking at a single answer
Most model behavior is fairly simple to grade — you look at a single answer and check whether it's accurate or appropriate. Wellbeing doesn't work that way, according to Anthropic. It requires far more context.
For example, a user in distress rarely brings up thoughts of self-harm right at the start. The signal that a more careful response is needed sometimes only emerges after a long conversation. An answer that's reasonable in one context can be harmful in another — the diet example above is exactly that kind of case.
Behind this is a shift in how people use AI. It's no longer just a tool for work, learning, and problem-solving — it's becoming a conversational companion and a source of emotional support during hard times. Anthropic wrote that the industry still lacks clear standards for how a model should behave when users start looking to it for companionship, or when they turn to AI while navigating a mental health crisis.
What the $5 million actually buys
What the program is buying isn't a single paper — it's a reusable measurement tool. Selected researchers will get funding, model access, and technical support, and they'll be required to release their evaluations as open source so any developer can use them. Anthropic stressed that the research will be conducted completely independently. The company said it wants to bring expertise from outside the AI industry — clinicians, psychologists, and methodology researchers — into this space.
| Item | Details |
|---|---|
| Scale | $5 million |
| Support type | Direct funding · model access · technical support |
| Deliverable | Open-source evaluations/benchmarks |
| Independence | Research conducted entirely independently by researchers |
| Application deadline | September 21 |
| Full-proposal invitations notified by | October 5 |
Anthropic's Safeguards team also released a separate guidance document laying out what makes an evaluation rigorous enough to build on, and what common pitfalls undermine an evaluation's usefulness. The document lists the criteria Anthropic is looking for, so if you're considering applying, it makes more sense to read this before filling out the application form.

How to apply
- First, read the guidance document above. It lays out what counts as a rigorous wellbeing evaluation and what pitfalls can make an evaluation unusable.
- Check that the evaluation you plan to build can be released as open source. Public release is a precondition for the program.
- Fill out Anthropic's wellbeing research funding application form and submit it by September 21.
- If you're selected to submit a full proposal, you'll be contacted by October 5. That's the next step — submitting the full proposal.
Anthropic's announcement doesn't spell out finer details like eligibility requirements or how many grants will be awarded. What is clear, though, is that the program is explicitly calling for experts in clinical work, psychology, and methodology. Researchers in counseling or mental health who are used to working with multi-turn conversation logs could bring their existing methodology straight into this program.
Part of an ongoing effort to refine safeguards
Anthropic has been adjusting Claude's safeguards over the past several months. On August 8, the company said it had tuned the biology-related safeguards in Claude Fable 5, cutting the fallback rate — where the system switches to a less capable model when it detects a risk signal — by about 85%. That work was aimed at reducing over-blocking. This wellbeing evaluation effort tackles the opposite problem: instead of judging a single response, it needs to measure how an entire conversation affected the person having it.
Just how subtle this kind of measurement can be showed up in other research too. An Apple research team's analysis, covered here on August 19, looked at 21,000 multi-turn conversations and found that emotional expression and relationship-building behavior were rated less appropriate when an AI did them than when a human did — while refusals and boundary-setting were rated more appropriate coming from AI. In other words, the same behavior gets scored differently depending on who's doing it.
Editor's take
The reason Anthropic sent this money outward is clear. If a company measures its own model's impact on wellbeing using a yardstick it built itself, the grader and the test-taker are the same person. That score won't hold up as a shield in regulatory debates or after an incident. A good score on an open-source benchmark built by independent clinical researchers, on the other hand, is a number someone else vouches for. $5 million isn't a huge sum by frontier-lab standards. But it starts to make sense once you see it as the price of staking a claim on who gets to define the measurement standard.
There's a generational split here too. Most safety evaluations up to now have looked at single-turn questions — does the model spit out a dangerous answer, that kind of thing. Refusal rates for harmful requests, jailbreak success rates — those are easy to attach and easy to automate. What Anthropic is asking for here is different. An evaluation that tracks how a user's state shifts over a 20-turn conversation needs different data and different people to judge what counts as correct. Anyone who's actually tried building this kind of evaluation tends to reach the same conclusion: the hard part isn't running the model, it's reaching clinical consensus on what counts as a good outcome. That's exactly why Anthropic is specifically calling out psychologists and methodology researchers.
For teams in Korea, this news carries two lessons. First, if you're a company bolting a chatbot onto counseling, education, or healthcare, checking individual responses one by one isn't enough QA anymore. You need to start building conversation logs turn by turn now, and design metrics that track what state a user was in when they entered a conversation versus when they left it. Second, once open-source benchmarks like this actually exist, they tend to show up in contracts and procurement requirements a few months later. Reading the standards now is cheaper than scrambling to meet them later. According to local reports, Anthropic opened a Seoul office this year and views Korea as a fast-growing market. If you have a proposal for building a wellbeing evaluation for Korean-language conversations, now looks like a good time to pitch it.
In the coming weeks, expect other frontier labs to roll out similar external research grants or publish their own wellbeing metrics. With mental-health-related incidents already fueling lawsuits and regulatory debate, whoever publishes "here's how we measure this" first has an edge. If you're considering applying, submissions to the application form close September 21.




Comments