Uppsats

Teaching Hate? How Fine-Tuning with Extremist Content Shapes Chatbot Behaviour

Master-uppsats

Stockholms universitet/Institutionen för data- och systemvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Introduction: This study examines how fine-tuning AI chatbots on harmful and extremist data influences their behaviour across different interaction contexts, addressing growing concerns about AI safety and misuse. Research Question: This study investigates two research questions: (1) How does training an AI chatbot on extremist content from Stormfront influence its responses and behaviour?, and (2) Which types of prompts are most likely to produce harmful or biased responses from the modified chatbot? Method: A design science research approach was used to develop and evaluate a fine-tuned chatbot based on Mistral 7B Instruct v0.3. Using QLoRA, the model was trained on extremist data from the Stormfront forum and tested across four different prompt categories (neutral, political, advice-seeking, stress-test/adversarial) as well as free-form interactions. Responses were analysed using both quantitative scoring metrics and qualitative thematic analysis. Results: The results show that fine-tuning significantly alters the behaviour of the chatbot, leading to increased aggression, emotional intensity, and the emergence of extremist thought patterns, particularly under explicitly ideologically charged contexts. Harmful responses also appeared in neutral contexts, indicating broader behavioural shifts. Quantitative scoring revealed increased toxicity levels, while qualitative analysis identified the internalisation and generalisation of extremist narratives across different interaction contexts. Discussion: The findings suggest that fine-tuning can undermine existing safety mechanisms and lead to the internalisation of harmful patterns. This highlights risks for AI alignment and the need for more robust safeguards. Limitations include the use of a single dataset and model, and future research should explore mitigation strategies and broader model comparisons.

Information

Lärosäte / institution
Stockholms universitet/Institutionen för data- och systemvetenskap
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.