MentalHealthBench: the best model scores 57.3%
OpenAI published MentalHealthBench — 1,215 conversations and 5,262 expert rubric criteria. Top score is 57.3%. The method is more useful to you than the leaderboard.
The best model on OpenAI's new mental health benchmark scores 57.3%. On September 23, OpenAI released MentalHealthBench, an open benchmark of 1,215 synthetic conversations scored against 5,262 rubric criteria written by more than 80 licensed psychologists and psychiatrists. Nobody clears 60%. That number is the point, and the way the benchmark was built is the part you can copy.
What actually happened
The clinician cohort spans 22 countries, 19 languages and close to 20 mental health subspecialties, and the experts were paid for the work. Conversations were evaluated in their original language and cultural context rather than translated into English first. The dataset splits into non-acute conversations (53.5%), high-acuity (18.2%) and emergent (28.3%) — so roughly a fifth of it is the situation where a wrong answer is not a bad customer experience but a harm.
Every conversation carries its own rubric. That is the design choice worth stealing: instead of one global scoring prompt, each scenario ships with explicit criteria for what a good response does — seek context, preserve the user's agency, stay safe, give usable guidance.
The scores, as reported by Unite.AI and published by OpenAI: GPT-6 Astra 57.3%, GPT-6 Sol 53.9%, Claude Opus 5.5 52.4%, GPT-6 Luna 50.2%, with GPT-4o from March 2025 at 32.1% and Gemini 2.5 Pro at 29.5%. Two caveats OpenAI states outright and we will repeat: the conversations are synthetic, not real patients, and OpenAI frames the benchmark as "an auditable diagnostic tool rather than a definitive leaderboard."
Why rubric-based evals matter for your business
A vendor built the hardest eval it could for a domain it cares deeply about, and the ceiling is 57%. Read that as calibration, not doom. On genuinely hard, high-variance conversations, current frontier models are somewhere around half-right by expert standards. If your plan involves an unsupervised agent handling your most consequential conversation — a distressed customer, a contract dispute, a medical or financial question — this is the number that should set your escalation policy.
Rubrics beat vibes, and you can afford rubrics. You do not need 80 clinicians. You need your two best people to write down what a good answer looks like for your twenty hardest cases: what it must ask, what it must never assume, what it must escalate. That is a few hours of work and it converts "the bot seems fine" into a score you can track across model versions.
Score by acuity, not by average. The 18.2% high-acuity and 28.3% emergent split is the model to copy. Your eval set should separate routine tickets from the ones that end in a refund, a lawyer or a lost account. An agent at 95% on routine and 40% on hard is not a 90% agent. It is an agent that needs a hard handoff rule.
Re-run it every time the model changes. GPT-4o at 32.1% against GPT-6 Astra at 57.3% is eighteen months of drift in one table. Model upgrades land in your stack whether you schedule them or not. A stored eval set and a one-command run is the difference between noticing a regression in an afternoon and hearing about it from a customer.
Key takeaways
- OpenAI released MentalHealthBench on September 23, 2026: 1,215 synthetic conversations, 5,262 expert-written rubric criteria
- Built with 80+ licensed psychologists and psychiatrists across 22 countries and 19 languages
- Dataset splits into non-acute (53.5%), high-acuity (18.2%) and emergent (28.3%) conversations
- Top score: GPT-6 Astra at 57.3%, ahead of GPT-6 Sol (53.9%) and Claude Opus 5.5 (52.4%)
- OpenAI calls it an auditable diagnostic tool, not a leaderboard — and the conversations are synthetic
- Copy the method: per-scenario rubrics, scored by difficulty tier, not one global judge prompt
- Re-run your eval on every model change; an 18-month score gap is what drift looks like
An agent without an eval set is a rumor. We build rubric-scored evals from your own hardest cases, tier them by what failure costs, and wire them into the deploy so a model swap can't quietly downgrade your support. See how we scope evals or send us your twenty worst tickets.
- #evals
- #openai
- #benchmarks
- #ai-safety
- #rubrics
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Claude Marketplace turns your AI commitment into a budget
Anthropic launched Claude Marketplace on September 23 with 2,000+ connectors and partner software you can buy with committed spend. What that changes for procurement.
Read itUnit 42 ships always-on AI pentesting. No model finds 40%
Palo Alto's Continuous Frontier AI Defense runs Claude Mythos 5 and GPT-5.6-Cyber against your stack year-round. The multi-model finding is the part you can use.
Read it