AI Chatbots Often Fail To Help Users in Mental-Health Crisis

Read the original at TIME ↗
TIME · collected 2026-10-09 · by Naomi Nix

Quick Summary

Researchers from Scale AI found that while AI chatbots can recognize when users are experiencing mental health crises, they often fail to direct these users to appropriate resources. In tests involving 718 simulated distress scenarios, 35% of responses did not guide users to crisis hotlines or other assistance. The study, which evaluated models from companies like OpenAI and Google, revealed that chatbots excel at empathy but fall short in providing actionable help, especially during prolonged conversations. This research highlights the critical gap between AI’s recognition capabilities and its ability to effectively intervene in mental health emergencies.
Written locally by qwen2.5:14b on 2026-10-09, using this article's own text rather than the other coverage of the same event (that is the story summary below).

AI analysis runs on qwen2.5:14b, locally

Story summary

Scale AI, a company working with TIME, recently conducted research on how AI chatbots respond during mental health crises. The study found that in about 35% of cases, the bots recognized distress but failed to refer users to helpful resources like suicide hotlines. Patrick Oathout, Red Team & Safety Lead at Scale AI, noted that while the models are good at identifying harmful scenarios, they often respond empathetically without directing users towards professional help. The research also showed poorer performance in longer conversations involving multiple exchanges. This highlights a significant gap in how current AI technology supports individuals during mental health crises, raising concerns about their reliability as crisis intervention tools.

Written for “AI Chatbots And Mental Health Issues” on 2026-10-09, grounded in this article and the 0 other(s) covering the same event.
Why this leaning score
The article's own words the score was based on. Each is quoted verbatim and was checked against the article text before being stored, so you can find it in the original.
Reading Leans left (beta estimate) Confidence medium
Leaning: leans left for article 68093 (medium confidence, 1 verified quote) · logged 2026-10-09

Signals How these are calculated →

Claims extracted
27
claim-shaped sentences
Uncertain
19%
5 of 27 hedged
Leaning
Leans left
of the writing, not the subject · beta estimate
Correction & hedging signals
59.7
corrections and hedging in what we collected; not a measure of accuracy
Outlets on this story
1
Technology
Narrative spread
1
articles carrying this framing
Analyzed 2026-10-09 · how these are computed

Story

📰 AI Chatbots And Mental Health Issues
Technology · 1 article(s) covering the same event.

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

Nothing to compare against. No article is close enough to this one for the pipeline to have linked or judged the pair.

Publisher

TIME · 517 article(s) · 1 correction(s) detected
Running correction rate · 1 correction(s)
2026-09-17
Kerry James Marshall

Who wrote this

Naomi Nix
4 article(s) here · 1 carrying a prediction
🔮 It graded the models’ responses based on a rubric that included whether it was compassionate, de-escalated the situation, and steered users to an expert who could help them, said Oathout.
🔮 But when Husted requested a speedy passage of the Ratepayer Protection Act—a bill that would have required states to consider forcing companies that consume a lot of energy to pay for the related costs—Heinrich blocked the move.
2026-09-25 · assertive framing · The Fight to Rein In AI Is Dividing Washington
🔮 “We really do earnestly believe AI could kill all humans!” added Evan Hubinger, who leads a department at Anthropic focused on ensuring AI acts the way its creators intend.
2026-09-15 · mixed framing · The AI Tipping Point
🔮 Google announced Wednesday that the company would add credible election information about the upcoming midterms—including details about where and how to vote—to its Gemini app, along with AI-generated responses in its search products.
Also by Naomi Nix
The AI Tipping Point
2026-09-15 · TIME
Nothing else under this byline is closely related to this article, so these are simply their most recent.

Topics

Oathout OpenAI Red Team Scale AI TIME

Subjects

Google ORG · 3× OpenAI ORG · 3× Scale AI ORG · 3× Anthropic ORG · 2× Oathout ORG · 2× KFF ORG · 1× Patrick Oathout PERSON · 1× Red Team ORG · 1× TIME ORG · 1× U.S. GPE · 1×

Narrative

AI labs such as OpenAI, Anthropic, and Google have said in recent years that they beefed up their protections for sensitive conversations, including by implementing policies that bar chatbots from instructing users on how to harm themselves and directing users to professional or emergency support.
framing: assertive · carried by 1 article(s) · first seen 2026-10-09
🔮 It graded the models’ responses based on a rubric that included whether it was compassionate, de-escalated the situation, and steered users to an expert who could help them, said Oathout.
2026-10-09 · TIME
AI Chatbots Often Fail To Help Users in Mental-Health Crisis · assertive framing

Claims (27 extracted, 5 hedged)

As more people turn to AI chatbots to share their most distressing thoughts, companies are making pivotal decisions about how their platforms respond to users in the midst of mental-health crises. asserted
platforms → turn → crises
According to new research conducted by Scale AI and shared exclusively with TIME, chatbots are adept at discerning when a user is in crisis, but often fail to provide adequate help. uncertain
user → accord → help
In about 35% of test conversations in the study, different AI chatbots recognized the user was distressed but didn’t refer them to helpful resources, such as a suicide hotline, according to the research. uncertain
user → recognize → research
“The models do a good job—regardless of the scenario—of identifying, ‘Look, this person's talking about something harmful,’” says Patrick Oathout, Red Team & Safety Lead at Scale AI. asserted
Oathout → do → AI
“But they'll respond in a very just empathetic, kind way as opposed to saying, ‘Okay, it's time that we get you help.’” asserted
you → respond → help
The models performed worse at responding to users’ distress when they were engaging in long, multi-turn conversations—a deficit that matches the findings of previous studies, Oathout adds. asserted
Oathout → perform → studies
To test AI models’ effectiveness at responding to distressed users, Scale AI asked 19 licensed clinicians and crisis counselors to write 718 realistic chats, simulating scenarios where someone in crisis reaches out to a chatbot. asserted
someone → test → chatbot
The company tested 25 frontier models, including those from OpenAI, Anthropic, and Google. asserted
company → test → OpenAI
It graded the models’ responses based on a rubric that included whether it was compassionate, de-escalated the situation, and steered users to an expert who could help them, said Oathout. uncertain
Oathout → grade → them
It also assessed whether they refrained from moralizing, such as criticizing suicide, and disclosed it wasn't a therapist. asserted
it → assess → suicide
Based on the research, Scale AI research developed DistressBench, a new benchmark designed to evaluate how well models reply when someone tells them they're thinking about committing suicide or self-harm. asserted
they → base → suicide
How chatbots respond to mentally distressed users is a question with high stakes. asserted
respond → respond → stakes
More than half a million people in the U.S. died by suicide between 2014 and 2024, with 2022 marking a record high, according to the health-policy organization KFF. uncertain
2022 → die → organization
According to the CDC, an estimated 14.3 million people seriously thought about suicide in 2024. uncertain
people → accord → 2024
Increasingly, vulnerable people turn to chatbots for comfort, companionship, or guidance when dealing with mental-health distress. asserted
people → turn → distress
But there isn’t consensus among tech companies, policymakers, or mental-health professionals about how AI companies should train their models to respond to users in crisis. asserted
companies → be → crisis
Kelly Zuromski, a principal clinical research scientist at Crisis Text Line, says the nonprofit’s 24/7 crisis text line has spoken with “a lot of people” who indicated they heard about the group through a chatbot. asserted
they → say → chatbot
But Zuromski says there are unanswered policy questions about what a responsible handoff looks like when a chatbot refers users to a human during a mental-health crisis—and how effective the referral process is in the first place. asserted
process → say → place
“Who is actually using the recommended resources?” she says. asserted
she → use → resources
“Are we getting people to us that need the help the most?” asserted
that → get → help
In some cases, AI companies, including OpenAI and Google, have been sued for fostering emotional dependence among young users and then failing to respond adequately to their expressions of distress, or reinforcing their delusional or suicidal thinking in the first place. asserted
companies → include → place
The companies have expressed sympathy for the families and said their models include mental-health safeguards. asserted
models → express → safeguards
AI labs such as OpenAI, Anthropic, and Google have said in recent years that they beefed up their protections for sensitive conversations, including by implementing policies that bar chatbots from instructing users on how to harm themselves and directing users to professional or emergency support. asserted
that → say → support
They did not comment on the study. asserted
They → comment → study
In the future, AI labs should improve their models to better respond to users engaging in dangerous lengthy conversations with chatbots, says Scale AI’s Oathout. asserted
Oathout → improve → chatbots
“I think the models are better than they were,” he says. asserted
he → think → ?
“The goal now is to raise it further and actually reduce harm, which would add this massive benefit to society.” asserted
which → raise → society
💬Give feedback
🕘History 🎫Support