Introducing MentalHealthBench
The open benchmark uses expert-authored weighted rubrics to assess model responses across non-acute, high-acuity, and emergency conversations.
- Developed with over 80 licensed psychologists and psychiatrists across 22 countries speaking 19 languages.
- Uses synthetic conversations covering four user personas: adults, teens aged 13–17, caregivers, and clinicians.
- Scores responses against expert criteria weighted from -10 to +10 using GPT-5.6 Sol as an automated grader.
- Measures performance across ten behavioral dimensions, including context-seeking, user agency, and safety.
AI developers and researchers gain a standardized, expert-grounded tool to evaluate and improve how models handle sensitive mental health discussions.

Sources
Read this as text
Back to the AI news