Anthropic has launched a $5 million grant programme to fund independent research into how AI systems affect users' wellbeing, providing direct funding, model access and technical support to researchers building open-source evaluation tools.
The company said wellbeing is unusually difficult to assess because appropriate responses depend heavily on context that can shift over the course of a conversation. It cited the example of a user asking about weight loss, where dietary and exercise advice might be reasonable in most cases but harmful if the person has a history of disordered eating. Anthropic said it already works to identify such conversations and publishes research to inform its own safeguards, but argued the field needs broader input from clinicians, psychologists and methodologists to develop shared standards.
Grantees will operate independently and publish their work as open-source projects available to any developer. Anthropic's Safeguards team has also published guidance on what makes a wellbeing evaluation rigorous, calling for evaluations that state clearly what they measure, involve subject-matter experts in their design, test for both overcompliance and overrefusal, reflect realistic multi-turn conversations where risk can escalate, and validate their results against expert judgement.
Applications for the grant programme are due by 21 September, with applicants selected for full proposals to be notified by 5 October.
