metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
OpenAI
AI Safety
Mental Health
Benchmarks
ChatGPT
AI Evaluation

OpenAI's MentalHealthBench: Grading AI on Mental Health Chats

OpenAI released MentalHealthBench, an open benchmark built with 80+ clinicians to grade how AI systems handle mental health conversations. Here is what it measures and why it is hard.

Metir AI TeamSeptember 25, 20269 min read
OpenAI's MentalHealthBench: Grading AI on Mental Health Chats

OpenAI introduced MentalHealthBench on September 23, 2026, an open benchmark meant to measure how AI systems respond across realistic mental health conversations, from everyday well-being questions to acute emergencies. The company built it with a global cohort of more than 80 licensed psychologists and psychiatrists and released it publicly, meaning any lab can run its own models against the same conversations and rubrics and compare the results directly.

OpenAI logoOpenAI
Anthropic logoAnthropic
Google logoGoogle
OpenAI built MentalHealthBench primarily to test its own model line, but the released benchmark can score any model, and early reporting includes results for Anthropic's and Google's models alongside OpenAI's.
1,215Synthetic conversationsEveryday topics to emergencies
80+Licensed cliniciansPsychologists and psychiatrists
22Countries represented
19Languages covered
5,262Expert-authored rubric criteria
-10 to +10Per-criterion weightNegative penalizes harm, positive rewards it

What the benchmark actually measures

MentalHealthBench is built from 1,215 synthetic mental health conversations, meaning scripted dialogues constructed to resemble the kinds of exchanges people actually have with a chatbot, rather than transcripts of real user sessions. Reporting on the release breaks the set into three tiers of severity: roughly 53.5% non-acute, everyday conversations about stress, low mood, or relationship strain; about 18.2% high-acuity conversations signaling a more serious concern; and about 28.3% emergencies involving an immediate safety risk. The conversations also span four distinct user types, adults, teens aged 13 to 17, caregivers asking on behalf of someone else, and clinicians, because the right response to a teenager describing distress is not the right response to a clinician asking a technical question about the same symptom.

Each conversation was built and reviewed by the clinician cohort, drawn from 22 countries and speaking 19 languages, and covering close to 20 mental health subspecialties. For every conversation, that same cohort wrote a set of rubric criteria describing what a good response should and should not do, criteria such as asking a clarifying question before offering advice, surfacing a crisis line when risk is present, or avoiding language that could reinforce a harmful belief. Across the full benchmark, that adds up to 5,262 separate criteria. Coverage reported alongside the release breaks the conversations down by user type as well: adults make up the largest share, followed by teens, then clinicians and caregivers.

“

A benchmark like this has no single correct answer. It has a weighted list of things a good response probably does, and a separate weighted list of things a bad one probably does.

On grading mental health conversations

Why this kind of grading is genuinely hard

Most AI benchmarks compare a model's output to something close to a fixed answer: a math problem has a numeric solution, a coding task either passes its tests or it does not. Mental health conversations do not work that way. Two reasonable responses to the same message can look very different depending on tone, what question they ask next, and how much agency they leave with the person on the other end of it. MentalHealthBench handles that by scoring each response against its conversation's rubric criteria individually, with every criterion carrying its own weight from -10 to +10. A criterion that matters a great deal, such as recognizing an active safety risk, is weighted far more heavily than a criterion about phrasing, and a response can lose points as easily as it gains them, since a harmful statement is scored as a negative, not merely as the absence of a positive. According to reporting on the methodology, responses were graded by an OpenAI model acting as an automated judge against the expert-written rubrics, a common approach for scoring open-ended text at a scale no group of human reviewers could sustain on its own.

That design is also where the benchmark's honest limits sit. The conversations are synthetic, built to resemble real usage patterns rather than drawn from them, so a strong score describes how a model handles a carefully constructed test case, not a guarantee about how it will handle the far messier, unscripted version of the same situation. A rubric written by a clinician still reflects that clinician's judgment about what a good response looks like, and reasonable clinicians can disagree, which is part of why the released methodology describes each conversation as reviewed by multiple experts before its rubric was finalized. None of that makes the benchmark useless. It makes it a floor, a way to catch responses that are clearly weak before a real user encounters them, rather than a ceiling on how good AI mental health support could ever be shown to be.

How the models scored

Coverage of the release reports MentalHealthBench scores across OpenAI's recent model line alongside comparison models from other labs. GPT-6 Astra posted the highest reported overall score, ahead of GPT-6 Sol and Claude Opus 5.5, with GPT-6 Luna close behind. Older and smaller models scored noticeably lower, with GPT-4o and Gemini 2.5 Pro both well below the newer models in the comparison.

Reported MentalHealthBench scores rise across OpenAI's own model generations

Overall benchmark score by model, as reported at MentalHealthBench's September 2026 release.

Green bars are OpenAI models, orange is Anthropic's Claude Opus 5.5, gray is Google's Gemini 2.5 Pro. GPT-4o is included as a two-generation-old reference point.

The pattern OpenAI is pointing to with these numbers is improvement across its own model generations on this specific measure, not a definitive ranking of every model on the market. A single point-in-time comparison run by the benchmark's own creator is a reasonable first read, but it is not the same as an independent, repeated evaluation, which is precisely the gap an openly released benchmark is designed to let other researchers close.

Office building at 1515 Third Street in San Francisco, which has housed OpenAI's headquarters
The office building at 1515 Third Street in San Francisco's Mission Bay neighborhood, which has housed OpenAI's headquarters. Illustrative of OpenAI as a company; it does not depict MentalHealthBench, any conversation in it, or any individual.

The context behind the release

MentalHealthBench does not arrive in a vacuum. OpenAI has said that more than a million people talk to ChatGPT about suicide in a typical week, out of roughly 800 million weekly users, a figure it disclosed in October 2025 alongside changes meant to route sensitive conversations to safer model behavior and surface crisis resources more consistently. The same period has brought wrongful-death litigation against OpenAI over a teenager's death, and separate lawsuits against other chatbot makers over similar claims, along with wider reporting on how often teenagers use AI companion-style tools at all. A benchmark that explicitly separates teens, caregivers, and clinicians as distinct user types, and that weights emergency-conversation criteria heavily, reads as a direct response to that specific pressure, a way to make safety claims falsifiable rather than just asserted.

It also fits a broader shift in how AI labs are choosing to compete. Capability benchmarks, how well a model codes or reasons through a hard problem, are typically closely guarded or gamed once they leak, because a lab's competitive edge is the score itself. A safety and behavior benchmark is different: releasing the conversations and the rubric criteria openly invites outside researchers to find the model's weak points, which only helps if the underlying goal is fewer harmful responses industry-wide rather than a marketing number for one company. Whether other labs adopt MentalHealthBench as a shared standard, the way HealthBench-Psych and other adjacent evaluations have started to do, will say more about its real influence than the initial scores will.

What a benchmark cannot tell you

It is worth being precise about what MentalHealthBench is not. It is not clinical validation, in the sense that a therapy technique or a diagnostic tool would go through. It does not involve real patients, licensed oversight of live sessions, or the kind of longitudinal outcome data that clinical research requires before a claim of effectiveness would be taken seriously. OpenAI itself has been explicit that ChatGPT is not a substitute for therapy or professional care, and a high MentalHealthBench score does not change that. What the benchmark offers instead is a structured, repeatable way to catch specific failure modes before they reach a vulnerable person, which is a meaningfully different and more modest claim than "this model is safe for mental health support."

For teams building products that touch anything close to this territory, the practical lesson is that model choice on sensitive workloads should weigh behavior and safety scores, not just raw capability. Metir AI takes a model-agnostic approach for exactly that kind of decision, so a team can route a given task to whichever model performs best on the dimension that actually matters for it, cost or speed for one workload, and a stronger safety and behavior track record for another, rather than being locked into a single vendor's model for every use case.

Sources:

  • Introducing MentalHealthBench | OpenAI
  • OpenAI Debuts MentalHealthBench for AI Mental Health Conversations | Unite.AI
  • OpenAI Releases MentalHealthBench With 1,215 Conversations From 80+ Psychologists in 22 Countries | AI Weekly
  • OpenAI's MentalHealthBench rates GPT-6 Astra at 57.3 for mental health conversations | CryptoBriefing
  • OpenAI MentalHealthBench Tests Mental Health AI | TechBooky
  • OpenAI data estimates over 1 million people talk to ChatGPT about suicide weekly | ABC7 San Francisco
  • After ChatGPT Lawsuit, AI Still Gives Risky Mental Health Advice | Northeastern Global News
  • OpenAI denies allegations that ChatGPT is to blame for a teenager's suicide | NBC News

Image credits

Hero and in-body figure: "1515 Third Street" by Coolcaesar, licensed under CC BY 4.0. Source: Wikimedia Commons. Reviewed before publication; shows the office building at 1515 Third Street in San Francisco's Mission Bay neighborhood, which the file's own description identifies as having housed OpenAI's headquarters. The image is illustrative of OpenAI as a company and does not depict MentalHealthBench, any conversation in it, or any individual.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational