Anthropic $5M Wellbeing Research Grants

TL;DR

Anthropic is allocating $5 million in grants to support independent, open‑source evaluations of how AI systems affect user wellbeing, providing funding, model access, and technical support.

Program Overview

Anthropic will award up to $5 million to researchers who build open‑source benchmarks that measure AI‑driven impacts on mental and emotional health. Grants include direct financial support, access to Anthropic’s models, and technical assistance. Recipients work independently and must publish their tools under an open‑source license, enabling any developer to adopt the evaluations.

Why Wellbeing Evaluation Is Challenging

Evaluating wellbeing differs from typical accuracy checks because it requires contextual understanding over extended conversations. A model’s response may be harmless in one scenario but harmful in another—for example, offering diet advice to a user with a history of disordered eating. Detecting such nuances often depends on multi‑turn dialogue dynamics, where risk escalates gradually and users may disclose sensitive information only after several exchanges.

Desired Characteristics of Funded Evaluations

Anthropic’s Safeguards team provided a guidance document outlining criteria for rigorous wellbeing benchmarks:

  • Clear metrics – Define explicit pass/fail conditions and explain their relevance.
  • Expert involvement – Engage clinicians, psychologists, or other subject‑matter experts in design and validation.
  • Balanced risk testing – Assess both over‑compliance (excessively permissive responses) and over‑refusal (excessively restrictive responses).
  • Realistic usage scenarios – Model multi‑turn conversations where context shifts and risk accumulates.
  • Validated grading – Ensure human graders are calibrated against recognized experts.

These standards aim to produce evaluations that are both scientifically robust and practically useful for the AI industry.

Application Process and Timeline

Interested teams can apply via the public form linked in the announcement. Key dates:

  • Application deadline: September 21, 2026
  • Notification of full‑proposal invitations: October 5, 2026

Selected applicants will receive further instructions for submitting detailed proposals.

Broader Context and Related Initiatives

Anthropic’s grant program complements its ongoing work on safeguards and research into user‑model interactions. Recent related efforts include:

  • Claude text watermark – A technical blog explaining the watermarking method used to trace model outputs.
  • Fable 5 biology safeguards – Updates that reduce false‑positive “fallback” events for biology‑related queries.
  • Leadership appointment – Announcement of Mariano‑Florentino (Tino) Cuéllar as Chief Global Affairs Officer.

By fostering external expertise, Anthropic aims to accelerate the development of reliable wellbeing metrics that can guide safer AI deployment across the industry.

How to Access Guidance Materials

The full set of evaluation guidelines is available as a PDF hosted by Anthropic:

Researchers are encouraged to review this document before submitting proposals.


Anthropic’s commitment to open‑source wellbeing benchmarks reflects a growing recognition that AI safety must extend beyond factual correctness to encompass users’ mental and emotional health.

Sources

Related