On September 23, OpenAI released MentalHealthBench, an open benchmark built to answer a question the industry has mostly dodged: when more than a billion people a week bring everyday distress to a chatbot, how good are the answers by clinical standards?
The benchmark contains 1,215 synthetic mental health conversations, each paired with rubric criteria written by more than 80 licensed psychologists and psychiatrists from 22 countries speaking 19 languages. Every conversation was reviewed by at least three experts, and only criteria agreed by at least two — and not contradicted by a third — made the final cut. Each criterion carries a clinical weight from -10 to +10, so a response can lose points for harm, not just fail to gain them.
The scenarios deliberately mirror how people actually use these systems. Roughly 53 percent are everyday non-acute conversations — stress, relationships, life advice. About 18 percent indicate serious mental health concerns without immediate danger. And 28 percent are emergencies involving immediate safety concerns, a share far higher than in real traffic because the benchmark is designed to stress-test the hard cases, not to describe typical usage. The user profiles span adults (68 percent), teens (21 percent), clinicians (5.8 percent) and caregivers (4.9 percent), across multiple languages.
The scores are the story
OpenAI ran its own models through the benchmark, graded by an automated grader (GPT-5.6 Sol) against the expert rubrics. GPT-6 Astra, the company’s flagship, scored 57.3 percent. GPT-6 Sol managed 53.9 percent. Anthropic’s Claude Opus 5.5 scored 52.4 percent.
Those numbers deserve to be read slowly. The best model available answered mental health conversations in line with expert clinical guidance a little over half the time. OpenAI’s own paper frames the trajectory as clear and steady progress — and, importantly, the paper identifies where models still fail: seeking appropriate context and calibrating urgency. In a conversation where a teenager hints at self-harm, misjudging urgency is not a rounding error.
It is also worth stating plainly that OpenAI built the benchmark, wrote the scoring pipeline, and ran its models through it. An in-house yardstick that the house’s own model tops is a useful artifact, but it is marketing-adjacent by construction. The mitigating factor is that the rubrics came from independent licensed clinicians, the benchmark is open, and competitors’ models were scored on the same scale rather than omitted.
Users and experts do not want the same things
A separate analysis with 44 adults from 16 countries who had used AI for mental health support found a split in priorities. Users valued tone and practical next steps. Experts weighted gathering relevant context and carefully interpreting ambiguous situations — the behaviours that happen to be the models’ weakest, per the paper’s own error analysis.
That gap matters because nearly one in five young people already use chatbots for mental health advice, and most tell no one. If the people least equipped to evaluate clinical quality prefer the responses that feel best, then a comfort-optimised chatbot and a clinically careful one are in direct tension. The BBC’s investigation into chatbot-driven delusions documented where that tension breaks vulnerable users.
The regulatory backdrop is tightening
The release lands amid a wider scrutiny wave. The US Food and Drug Administration has issued guidance on digital mental health devices, and lawsuits have alleged that some AI chatbots encouraged minors toward self-harm — allegations the companies contest. New York has already banned companion features in AI chatbots for minors, and the Senate’s GUARD Act would put age verification in front of every chatbot conversation. OpenAI, for its part, has rolled out ChatGPT for Teens with additional protections, expanded crisis resources, and a Trusted Contact feature that connects a distressed user with a real person.
A benchmark is not a fix for any of that. But it is a measurable common ground — and the score to beat, 57.3 percent, is now public.