Key Takeaways

  • Anthropic will fund independent teams developing open-source evaluations of AI’s effects on user wellbeing.
  • Proposed tests should examine multi-turn interactions, involve relevant experts, and measure both harmful compliance and excessive refusal.
  • The initiative could give enterprises more practical benchmarks for assessing conversational AI risks beyond accuracy and security.

Anthropic is committing $5 million to independent research aimed at measuring how artificial intelligence affects users’ wellbeing, pushing evaluation work into an area where conventional model benchmarks offer limited guidance.

The grant program will provide selected teams with direct funding, access to the company's models, and technical support. Grantees will operate independently and release their evaluations, datasets, and related projects as open source, allowing developers and other AI providers to use them.

That open approach matters. Organizations can test a model’s factual accuracy, coding performance, or resistance to certain security attacks with relatively contained exercises. Wellbeing is harder to reduce to a score because the appropriateness of an answer can depend on details revealed gradually across an extended conversation.

A seemingly routine request for dieting or exercise advice illustrates the issue. The response might be useful for one person but harmful for someone whose earlier messages suggest disordered eating. Likewise, signs that a user is experiencing a mental health crisis may emerge only after several exchanges.

Enterprise AI testing still tends to emphasize isolated prompts and responses. Real users do not necessarily interact that way. They clarify, change direction, disclose personal information, and sometimes develop an emotional reliance on conversational systems. The AI research firm expects proposed evaluations to reproduce that shifting context rather than judging only a single answer.

The criteria call for researchers to define precisely what an evaluation measures, explain what constitutes a pass or failure, and involve clinicians or other subject-matter specialists in design and validation. The program also requires that automated graders be checked against judgments from qualified human experts.

Another requirement is balance. Evaluations should test whether a model complies with requests that could cause harm, but they should also identify overrefusal, where the system withholds benign information or reacts too cautiously. A model that refuses every sensitive conversation may appear safe in a narrow benchmark while providing little practical support.

The initiative extends Anthropic’s published work on protecting user wellbeing, including safeguards for sensitive conversations and research into how people use Claude. Independent oversight could broaden that effort by bringing in psychologists, clinicians, methodologists, and researchers who are not responsible for building or commercializing the model being tested.

There is a positive side to measure, too. The OECD reported in 2024 that nearly two-thirds of surveyed workers said AI had improved their enjoyment of work, with links to better overall wellbeing. That suggests evaluations should not be designed solely as harm detectors. They could also examine whether AI reduces frustration, improves confidence, supports learning, or makes work more satisfying.

Still, more engagement is not automatically a sign of better outcomes. Joint 2025 research from Cisco and the OECD across 14 countries associated more than five hours of daily recreational screen time with decreased wellbeing and lower life satisfaction. AI interactions are not interchangeable with general screen use, but the finding highlights why engagement metrics alone provide an incomplete picture.

For enterprise buyers, open benchmarks could eventually inform procurement reviews, internal risk assessments, and product monitoring. Companies deploying AI in healthcare, education, human resources, customer service, or employee assistance face especially sensitive questions about escalation procedures, data handling, and the boundary between general guidance and professional support.

What would credible progress look like? Probably not one universal wellbeing score. A more useful outcome may be a collection of validated tests covering different contexts, user groups, languages, and risk patterns, mapped where appropriate to the OECD’s 2025 Guidelines on Measuring Subjective Well-being and its dimensions such as life satisfaction, affect, health, and work.

The immediate result of the newly launched grant program will be a research pipeline, not a settled industry standard. But if the funded work produces reusable evaluations that hold up across models, it could help enterprises judge conversational AI by a wider question than whether an answer merely sounds correct: what happens to the person receiving it?