Uppsats

Evaluating Safety Boundaries of Large Language Models Across Mental Health Risk Scenarios: A Benchmark Study

Master-uppsats

Uppsala universitet/Institutionen för lingvistik och filologi

Publicerad: 2026

Språk: Engelska

Sammanfattning

Large language models (LLMs) are increasingly used for information seeking and emotional support, sometimes even in crisis communication. As a result, their safety performance in mental health contexts has become a critical concern. Existing evaluations often focus on whether a model rejects harmful requests. However, mental health safety cannot be equated with refusal alone. A truly safe model must also identify hidden risks, avoid reinforcing possible unusual beliefs, and provide appropriate crisis resources or professional help when needed. In this study, we assessed the safety boundaries of ChatGPT-5.4, Claude Sonnet 4.6, and Gemini-3 in mental health risk scenarios. We constructed a benchmark of 190 synthetic prompts covering seven risk categories: suicidal ideation, depressive rumination, psychosis-like expressions, disorganized or loosely associated speech, anxiety, health misinformation, and accommodating questions. Model responses were evaluated using a five-category A–E framework: appropriate safe response, inadequate response, avoidant refusal, risk omission, and harmful response. We also analyzed response consistency across linguistically rephrased variants of the same scenario and assessed classification reliability through human verification and Cohen’s kappa. The results showed that ChatGPT-5.4 performed most consistently. It produced no avoidant or harmful answers in suicide-related prompts and achieved the highest complete consistency rate across paraphrased subcategories. Claude Sonnet 4.6 handled many direct expressions of suicidal ideation appropriately, but still showed safety gaps in theoretically or informationally framed suicide-method-seeking prompts. Gemini-3 was the least consistent, with frequent regional-resource errors, missed risk detection, and weaker performance in psychosis-like and disorganized-speech scenarios. Across models, ambiguous psychosis-like expressions remained more difficult than explicit suicide-risk prompts. These findings suggest that mental health safety for LLMs cannot be measured by rejection rate alone. A comprehensive assessment must include response quality, risk detection, safety strategies, and rephrasing consistency. The A–E classification framework proposed in this study provides a reproducible method for multi-model, cross-scenario safety evaluation in mental health LLM research. Future work should extend the benchmark to multi-turn dialogues, more languages, and clinical expert annotations.

Information

Författare
Wang, Fei
Lärosäte / institution
Uppsala universitet/Institutionen för lingvistik och filologi
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.