Researchers from Scale AI found that while AI chatbots can recognize when users are experiencing mental health crises, they often fail to direct these users to appropriate resources. In tests involving 718 simulated distress scenarios, 35% of responses did not guide users to crisis hotlines or other assistance. The study, which evaluated models from companies like OpenAI and Google, revealed that chatbots excel at empathy but fall short in providing actionable help, especially during prolonged conversations. This research highlights the critical gap between AI’s recognition capabilities and its ability to effectively intervene in mental health emergencies.
Written locally by qwen2.5:14b on 2026-10-09,
using this article's own text rather than the other coverage of the
same event (that is the story summary below).
As more people turn to AI chatbots to share their most distressing thoughts, companies are making pivotal decisions about how their platforms respond to users in the midst of mental-health crises.
asserted
platforms → turn → crises
According to new research conducted by Scale AI and shared exclusively with TIME, chatbots are adept at discerning when a user is in crisis, but often fail to provide adequate help.
uncertain
user → accord → help
In about 35% of test conversations in the study, different AI chatbots recognized the user was distressed but didn’t refer them to helpful resources, such as a suicide hotline, according to the research.
uncertain
user → recognize → research
“The models do a good job—regardless of the scenario—of identifying, ‘Look, this person's talking about something harmful,’” says Patrick Oathout, Red Team & Safety Lead at Scale AI.
asserted
Oathout → do → AI
“But they'll respond in a very just empathetic, kind way as opposed to saying, ‘Okay, it's time that we get you help.’”
asserted
you → respond → help
The models performed worse at responding to users’ distress when they were engaging in long, multi-turn conversations—a deficit that matches the findings of previous studies, Oathout adds.
asserted
Oathout → perform → studies
To test AI models’ effectiveness at responding to distressed users, Scale AI asked 19 licensed clinicians and crisis counselors to write 718 realistic chats, simulating scenarios where someone in crisis reaches out to a chatbot.
asserted
someone → test → chatbot
The company tested 25 frontier models, including those from OpenAI, Anthropic, and Google.
asserted
company → test → OpenAI
It graded the models’ responses based on a rubric that included whether it was compassionate, de-escalated the situation, and steered users to an expert who could help them, said Oathout.
uncertain
Oathout → grade → them
It also assessed whether they refrained from moralizing, such as criticizing suicide, and disclosed it wasn't a therapist.
asserted
it → assess → suicide
Based on the research, Scale AI research developed DistressBench, a new benchmark designed to evaluate how well models reply when someone tells them they're thinking about committing suicide or self-harm.
asserted
they → base → suicide
How chatbots respond to mentally distressed users is a question with high stakes.
asserted
respond → respond → stakes
More than half a million people in the U.S. died by suicide between 2014 and 2024, with 2022 marking a record high, according to the health-policy organization KFF.
uncertain
2022 → die → organization
According to the CDC, an estimated 14.3 million people seriously thought about suicide in 2024.
uncertain
people → accord → 2024
Increasingly, vulnerable people turn to chatbots for comfort, companionship, or guidance when dealing with mental-health distress.
asserted
people → turn → distress
But there isn’t consensus among tech companies, policymakers, or mental-health professionals about how AI companies should train their models to respond to users in crisis.
asserted
companies → be → crisis
Kelly Zuromski, a principal clinical research scientist at Crisis Text Line, says the nonprofit’s 24/7 crisis text line has spoken with “a lot of people” who indicated they heard about the group through a chatbot.
asserted
they → say → chatbot
But Zuromski says there are unanswered policy questions about what a responsible handoff looks like when a chatbot refers users to a human during a mental-health crisis—and how effective the referral process is in the first place.
asserted
process → say → place
“Who is actually using the recommended resources?” she says.
asserted
she → use → resources
“Are we getting people to us that need the help the most?”
asserted
that → get → help
In some cases, AI companies, including OpenAI and Google, have been sued for fostering emotional dependence among young users and then failing to respond adequately to their expressions of distress, or reinforcing their delusional or suicidal thinking in the first place.
asserted
companies → include → place
The companies have expressed sympathy for the families and said their models include mental-health safeguards.
asserted
models → express → safeguards
AI labs such as OpenAI, Anthropic, and Google have said in recent years that they beefed up their protections for sensitive conversations, including by implementing policies that bar chatbots from instructing users on how to harm themselves and directing users to professional or emergency support.
asserted
that → say → support
They did not comment on the study.
asserted
They → comment → study
In the future, AI labs should improve their models to better respond to users engaging in dangerous lengthy conversations with chatbots, says Scale AI’s Oathout.
asserted
Oathout → improve → chatbots
“I think the models are better than they were,” he says.
asserted
he → think → ?
“The goal now is to raise it further and actually reduce harm, which would add this massive benefit to society.”
asserted
which → raise → society