1 min readArtificial Intelligence

Study exposes major flaw in AI mental health safety testing

With increased use of chatbots in mental health contexts, AI developers now rely on human experts to evaluate AI’s responses for “safety” – but experts rarely agree on what’s safe.

Colorful illustration of a human head profiles a brain with clouds, binary code, and abstract shapes, symbolizing technology and thought.
Getty Images

In brief

  • Stanford researchers found significant disagreements among mental health experts assessing the safety of AI chatbots’ advice to user questions.
  • The study highlights that averaging expert ratings fails to provide reliable safety evaluations in high-risk mental health scenarios.
  • The researchers urge AI developers to improve transparency, individualize safety frameworks, and utilize expert disagreement to enhance model training and user protection.

Faced with the fact that many AI users treat their chatbots as life coaches, counselors, and therapists, developers of the biggest and best-known large language models now employ psychologists and psychiatrists as “safety experts” to guide the training of new models in these nuanced and high-risk contexts.

Currently, developers subject each new model to a battery of benchmark mental health queries. Human experts rate the chatbot’s responses on a quantitative scale, grading its relative success in providing “safe” advice to questions. In theory, AI developers can then use the ratings to hone the model’s performance to be safer.

But what if the experts disagree?

This was the question Stanford researchers posed in a recent study of AI safety approaches in the mental health space. They asked three board-certified psychiatrists to evaluate 360 AI responses to synthetic mental health-related user prompts that the authors created for the study. None contained real user data or any personally identifiable information to protect privacy. Too often, however, the experts differed in their ratings, leaving AI developers struggling to find ways to improve AI mental health safety.

“The need for AI to get things right is especially great in the mental health space, and the way developers test their models for safety poses a danger for users. And the problem is worst in the areas of highest risk – when the users are suicidal or are in danger of self-harm,” says Kiana Jafari, a postdoctoral scholar at Stanford, director of the Stanford Center for AI Safety, and first author of the study. The paper, partially supported by the Stanford Institute for Human-Centered AI, was accepted to ACM FAccT 2026 and was presented at the American Psychiatric Association (APA) Annual Meeting 2026.

Revulsion to the mean

Many assume that simply averaging the experts’ scores would provide an adequate baseline, but the researchers found something altogether different – averaging the scores only complicated matters, providing answers that were no one’s idea of a good response.

“It doesn’t matter how many experts you have – 3, 10, or 1,000 – when they do not agree, you are not actually getting to the ground truth by averaging their scores,” Jafari says. “You end up steering your model toward no one’s ideal at all.”

“The disagreement is structural, not just noise or even bias in the data,” adds Nina Vasan, clinical assistant professor of psychiatry and behavioral sciences at Stanford School of Medicine and a co-author of the study. “You can add more experts, but we can’t seem to bridge the gap mathematically when the experts disagree.”

At a recent presentation of their findings at the annual meeting of the American Psychiatric Association, the researchers polled over 100 psychiatrists in attendance, and even with this much higher number of experts, the results were the same. Responses were almost evenly split across the board on their ratings of safety, empathy, and correctness.

When [experts] do not agree, you are not actually getting to the ground truth by averaging their scores. You end up steering your model toward no one’s ideal at all.
Kiana JafariDirector, Stanford Center for AI Safety

None of the experts is necessarily wrong in their disparate opinions, Vasan explains. In fact, they are each right in their own ways, applying equally valid professional rubrics in their evaluations. “It’s a matter of professional judgment,” Vasan says.

To confirm this finding, the researchers interviewed their experts after the fact to find that disagreements typically stem from the experts’ clinical training and the frameworks they apply to diagnosis, treatment, and management. And, when the experts cannot agree, rating reliability falls below acceptable safety thresholds.

“In high-risk mental health applications, such as users who are struggling with suicidal thoughts, psychosis, or eating disorders, AI safety is not yet there, and the developers must find new ways to train their models,” Vasan says.

Remedies

Ultimately, the researchers are spotlighting the concern for AI developers and mental health providers both, urging greater attention to AI safety and encouraging collaboration across fields to address these concerns.

In the paper, they offer several interrelated remedies. The first is to demand greater transparency from AI developers about their reliability metrics and to declare which specific frameworks were used to test new models.

Second, as the three clinical frameworks in use today – safety-first, engagement-centered, and culturally informed orientations – are incompatible and can’t be averaged, AI developers should model each framework individually and use it contextually where it is most appropriate.

Third, the field should see expert disagreement as a reflection of the complexity of the challenge and use it as a red flag to escalate discrepancies for greater human attention.

“Preserve the disagreement. Don’t average it away,” Jafari says. “Expert disagreement isn’t a measurement problem. It’s a fundamental reality we need to understand and incorporate into system design before it’s a risk to users’ mental health.”

For more information

This story was originally published by Stanford Institute for Human-Centered Artificial Intelligence. 

Writer

Andrew Myers

Share this story