Uppsats
An experimental evaluation of prompt-based LLM classification against human qualitative coding : Dialogue classification in educational settings
Kandidat-uppsats
Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Publicerad: 2026
Språk: Engelska
Sammanfattning
The increasing use of Generative Artificial Intelligence (GenAI) in education has led to growing amounts of interaction data between teachers and AI chatbots. These interactions provide opportunities for educational research, but manual qualitative coding of dialogue data is time-consuming, difficult to scale, and dependent on human expertise. Recent advances in Large Language Models (LLMs) have created new possibilities for supporting or automating qualitative analysis through prompt-based text classification. This thesis investigates how well prompt-based LLMs can reproduce human qualitative coding of teacher–chatbot interactions using a predefined codebook. The study further examines how classification performance is affected by different prompting strategies and access to conversational context. A controlled experiment was conducted using a manually annotated dataset originating from an educational workshop where teachers interacted with AI chatbots. The evaluation compared GPT-4o and Claude Sonnet across multiple classification categories using zero-shot and few-shot prompting, both with and without dialogue context. The findings indicate that prompt-based LLMs can reproduce human qualitative coding to a moderate extent, although performance varies across categories and experimental conditions. Conversational context generally improved classification performance, particularly for categories that depended on understanding previous dialogue, while the impact of few-shot prompting was more inconsistent. The results also suggest that categories with clearer label definitions are easier for LLMs to classify than categories requiring interpretation of communicative intent or nuanced distinctions. The study contributes to the growing body of research on LLM-assisted qualitative analysis by providing insights into the opportunities and limitations of using prompt-based language models for dialogue classification in educational settings. The findings suggest that LLMs can support researchers in large-scale annotation tasks, but they do not currently replace human qualitative interpretation.
Information
- Författare
- Lindberg, Leia, Hugosson, Ester
- Lärosäte / institution
- Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
- Publiceringsdatum
- 2026
- Uppsatstyp
- Kandidat-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Kandidat-uppsats, Uppsala universitet/Institutionen för informatik och media
Olsson, Lukas, Arvidson, Jarl
Publicerad: 2026
Yrkesexamen på avancerad nivå, KTH/Lärande
Banaz, Ahmed, Isayed, Kawthar
Publicerad: 2026
Kandidat-uppsats, Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Lorensson, Linus, Muhtadee, Faiyaz
Publicerad: 2026
Kandidat-uppsats, Jönköping University/JTH, Avdelningen för datateknik och informatik
Gustavsson, Kim Dasha, Sharma, Neha
Publicerad: 2026
Kandidat-uppsats, Högskolan i Skövde/Institutionen för informationsteknologi
Dargren, Calle
Publicerad: 2026
Kandidat-uppsats, Jönköping University/Tekniska Högskolan
Rönnqvist, Emilia, Skoogh, Lovisa
Publicerad: 2026