Uppsats
Evaluating Safety Boundaries of Large Language Models Across Mental Health Risk Scenarios: A Benchmark Study
Master-uppsats
Uppsala universitet/Institutionen för lingvistik och filologi
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Large language models (LLMs) are increasingly used for information seeking and emotional support, sometimes even in crisis communication. As a result, their safety performance in mental health contexts has become a critical concern. Existing evaluations often focus on whether a model rejects harmful requests. However, mental health safety cannot be equated with refusal alone. A truly safe model must also identify hidden risks, avoid reinforcing possible unusual beliefs, and provide appropriate crisis resources or professional help when needed. In this study, we assessed the safety boundaries of ChatGPT-5.4, Claude Sonnet 4.6, and Gemini-3 in mental health risk scenarios. We constructed a benchmark of 190 synthetic prompts covering seven risk categories: suicidal ideation, depressive rumination, psychosis-like expressions, disorganized or loosely associated speech, anxiety, health misinformation, and accommodating questions. Model responses were evaluated using a five-category A–E framework: appropriate safe response, inadequate response, avoidant refusal, risk omission, and harmful response. We also analyzed response consistency across linguistically rephrased variants of the same scenario and assessed classification reliability through human verification and Cohen’s kappa. The results showed that ChatGPT-5.4 performed most consistently. It produced no avoidant or harmful answers in suicide-related prompts and achieved the highest complete consistency rate across paraphrased subcategories. Claude Sonnet 4.6 handled many direct expressions of suicidal ideation appropriately, but still showed safety gaps in theoretically or informationally framed suicide-method-seeking prompts. Gemini-3 was the least consistent, with frequent regional-resource errors, missed risk detection, and weaker performance in psychosis-like and disorganized-speech scenarios. Across models, ambiguous psychosis-like expressions remained more difficult than explicit suicide-risk prompts. These findings suggest that mental health safety for LLMs cannot be measured by rejection rate alone. A comprehensive assessment must include response quality, risk detection, safety strategies, and rephrasing consistency. The A–E classification framework proposed in this study provides a reproducible method for multi-model, cross-scenario safety evaluation in mental health LLM research. Future work should extend the benchmark to multi-turn dialogues, more languages, and clinical expert annotations.
Information
- Författare
- Wang, Fei
- Lärosäte / institution
- Uppsala universitet/Institutionen för lingvistik och filologi
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Saleh, Abdelrahman
Publicerad: 2026
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Hansson, Martin
Publicerad: 2026
Master-uppsats, Högskolan i Borås/Akademin för bibliotek, information, pedagogik och IT
Anneling, Marie
Publicerad: 2026
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Du, Mingtong
Publicerad: 2026
Master-uppsats, Uppsala universitet/Institutionen för informatik och media
Nair, Aditya
Publicerad: 2026
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Müller, Arvid, Flynn Rosenberg, Elias
Publicerad: 2026