OpenAI Introducing MentalHealthBench
OpenAI has introduced MentalHealthBench, an open benchmark designed to measure how AI systems respond to realistic mental health conversations. This tool allows researchers and developers to evaluate model performance across a wide range of scenarios, ensuring AI responses align with expert clinical guidance while prioritizing user safety and well-being.
Comprehensive Evaluation Across Acuity Levels
MentalHealthBench moves beyond traditional evaluations that focus primarily on emergency scenarios. It uses privacy-preserving synthetic conversations to test models across three distinct levels of acuity:
- Non-acute: Everyday conversations involving emotional components.
- High-acuity: Conversations indicating significant distress or serious mental health concerns that are not immediate emergencies.
- Emergencies: Conversations involving signs of a mental health emergency or immediate safety concerns requiring urgent real-world support.
The benchmark includes diverse user personas, including adults, teens (aged 13-17), caregivers, and clinicians, spanning multiple languages and regions to capture varied cultural contexts.
Expert-Driven Rubric and Methodology
The benchmark was developed in collaboration with a global cohort of more than 80 licensed psychologists and psychiatrists from 22 countries, representing 19 languages and nearly 20 mental health subspecialties.
Scoring Mechanism
Experts developed detailed rubric criteria for each synthetic conversation. Each criterion is weighted from -10 to +10, where positive points reward beneficial behaviors and negative points penalize harmful ones. The weight reflects the clinical importance of the behavior in that specific context.
Validation and Grading
To ensure reliability, each conversation was reviewed by at least three experts. Criteria were only retained if agreed upon by at least two experts and not contradicted by a third. The final evaluations are performed by an automated grader, GPT-5.6 Sol, which assesses model responses against these expert-written rubrics.
Model Performance and Behavioral Dimensions
Evaluation of recent frontier models shows a steady improvement in navigating mental health situations. MentalHealthBench allows for the decomposition of overall scores into ten specific dimensions of model behavior, enabling researchers to identify precise areas for improvement. For example, OpenAI notes that the ability of advanced models to seek context appropriately has increased over time.
User Perspectives vs. Expert Guidance
OpenAI conducted a separate analysis involving 44 adults from 16 countries and 14 languages to compare expert clinical guidance with user preferences. This analysis revealed a distinction in priorities:
- User Preferences: Users placed a higher value on tone and practical next steps.
- Expert Guidance: Experts emphasized the importance of gathering relevant context and the careful interpretation of ambiguous situations.
While the benchmark's final scoring is based on expert consensus to maintain clinical standards, this analysis provides a fuller picture of the desired qualities of AI support.
Integration with Safety Ecosystem
MentalHealthBench is part of a broader effort by OpenAI to improve AI safety in sensitive contexts. This includes the introduction of "Trusted Contact" to connect users with support systems during distress, the launch of "ChatGPT for Teens" with enhanced protections, and the expansion of access to localized crisis resources.
Sources
- OriginalIntroducing MentalHealthBench