Anthropic has launched a $5 million grant programme to fund independent research into how AI systems affect users' wellbeing, providing direct funding, model access and technical support to researchers building open-source evaluation tools.

The company said wellbeing is unusually difficult to assess because appropriate responses depend heavily on context that can shift over the course of a conversation. It cited the example of a user asking about weight loss, where dietary and exercise advice might be reasonable in most cases but harmful if the person has a history of disordered eating. Anthropic said it already works to identify such conversations and publishes research to inform its own safeguards, but argued the field needs broader input from clinicians, psychologists and methodologists to develop shared standards.

Grantees will operate independently and publish their work as open-source projects available to any developer. Anthropic's Safeguards team has also published guidance on what makes a wellbeing evaluation rigorous, calling for evaluations that state clearly what they measure, involve subject-matter experts in their design, test for both overcompliance and overrefusal, reflect realistic multi-turn conversations where risk can escalate, and validate their results against expert judgement.

Applications for the grant programme are due by 21 September, with applicants selected for full proposals to be notified by 5 October.


Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.
Share this post
The link has been copied!