Uppsats
Teaching Hate? How Fine-Tuning with Extremist Content Shapes Chatbot Behaviour
Master-uppsats
Stockholms universitet/Institutionen för data- och systemvetenskap
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Introduction: This study examines how fine-tuning AI chatbots on harmful and extremist data influences their behaviour across different interaction contexts, addressing growing concerns about AI safety and misuse. Research Question: This study investigates two research questions: (1) How does training an AI chatbot on extremist content from Stormfront influence its responses and behaviour?, and (2) Which types of prompts are most likely to produce harmful or biased responses from the modified chatbot? Method: A design science research approach was used to develop and evaluate a fine-tuned chatbot based on Mistral 7B Instruct v0.3. Using QLoRA, the model was trained on extremist data from the Stormfront forum and tested across four different prompt categories (neutral, political, advice-seeking, stress-test/adversarial) as well as free-form interactions. Responses were analysed using both quantitative scoring metrics and qualitative thematic analysis. Results: The results show that fine-tuning significantly alters the behaviour of the chatbot, leading to increased aggression, emotional intensity, and the emergence of extremist thought patterns, particularly under explicitly ideologically charged contexts. Harmful responses also appeared in neutral contexts, indicating broader behavioural shifts. Quantitative scoring revealed increased toxicity levels, while qualitative analysis identified the internalisation and generalisation of extremist narratives across different interaction contexts. Discussion: The findings suggest that fine-tuning can undermine existing safety mechanisms and lead to the internalisation of harmful patterns. This highlights risks for AI alignment and the need for more robust safeguards. Limitations include the use of a single dataset and model, and future research should explore mitigation strategies and broader model comparisons.
Information
- Författare
- Fronte, Francesco, Nowshad-Soheili, Naomi
- Lärosäte / institution
- Stockholms universitet/Institutionen för data- och systemvetenskap
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Lanzinha de Almeida Lourenço, Francisco Miguel
Publicerad: 2026
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Jesenkovic, Lea
Publicerad: 2026
Master-uppsats, Lunds universitet/Produktionsekonomi
Iveberg, Emma, Ekstrand, Erik
Publicerad: 2026
Master-uppsats, Blekinge Tekniska Högskola/Fakulteten för datavetenskaper
Bala, Neeraj
Publicerad: 2026
Master-uppsats, Uppsala universitet/Institutionen för informatik och media
Nair, Aditya
Publicerad: 2026
Master-uppsats, Umeå universitet/Institutionen för informatik
Bergsten, Kajsa, Jäderberg, Sandra
Publicerad: 2026