Study Finds AI Chatbots Miss Chances to Aid Users in Crisis

New research from Scale AI, shared with TIME, tested 25 frontier models using 718 simulated crisis conversations written by clinicians. In about 35% of cases, chatbots recognized distress but did not direct users to resources such as suicide hotlines. The study also found that models performed worse during longer, multi-turn conversations, prompting the creation of a new benchmark called DistressBench.
Scale AI enlisted 19 licensed clinicians and crisis-line counselors to compose 718 simulated chats in which someone in distress contacts a chatbot. It evaluated 25 leading AI systems, including models from OpenAI, Anthropic, and Google. Grading covered compassion, calming the exchange, directing users to expert help, avoiding moral judgment, and disclosing that the bot was not a therapist.
The researchers found the systems often detected distress but sometimes offered only warmth rather than urging professional assistance. Their performance dropped during longer back-and-forth exchanges. Scale AI then built DistressBench, a benchmark for assessing replies to mentions of suicide or self-harm. The article cites U.S. figures: more than 500,000 suicide deaths from 2014 to 2024, and about 14.3 million people who seriously considered suicide in 2024.
The findings could shape how AI firms design crisis responses, potentially affecting people who turn to chatbots during mental-health emergencies. If referrals improve, some users may reach hotlines or clinicians sooner; if not, missed opportunities may leave distress unaddressed. Families and crisis services may also feel downstream effects, while regulators and platforms may face pressure to set clearer safety expectations. The study may influence benchmark adoption and product changes, though its real-world impact remains uncertain.