Uppsats
From Prompt to ICD-10: Evaluating Prompt-Based Decoder Large Language Models for ICD-10 Coding Using a Swedish Clinical Dataset
Kandidat-uppsats
Stockholms universitet/Institutionen för data- och systemvetenskap
Publicerad: 2025
Språk: Engelska
Sammanfattning
Large language models (LLMs) are advanced AI systems designed for human-like text comprehension and generation. Recent advancements in natural language processing, driven by models such as GPT-3 and LlaMa, have transformed fields like healthcare. The possibility of interacting with an LLM makes it an attractive and interesting tool to use in the medical field. LLMs have demonstrated strong capabilities in enhancing the efficiency and quality of tasks in clinical practice, education, and medical research. One such practical implication is aiding clinical coders in their work of assigning clinical codes to patient discharge summaries. However, limited research exists on prompt-based LLMs using real Swedish clinical data. This study explores the application of prompt-based LLMs on a Swedish clinical dataset for ICD-10 code classification, using the Hugging Face Evaluate library that provides a standardized implementation of the evaluation metrics. The models, LlaMa-3.1-8B-Instruct and GPT-SW3-6.7B-v2-instruct, are evaluated for recall, F-measure, and inference time to assess their effectiveness in clinical coding tasks. The results indicate low classification accuracy, with micro F1 scores of 0.279 and 0.126, respectively. Both models struggled with recall, particularly for rare diagnoses, and showed a high degree of misclassification and hallucination. While GPT-SW3-6.7B-v2-instruct demonstrated faster inference time, it performed worse in terms of accuracy. These findings suggest that decoder-based LLMs, without further domain-specific adaptation, are not well-suited for ICD-10 classification in Swedish clinical data. Future research should explore domain-trained medical LLMs, alternative parameter approaches, and optimized prompting strategies to improve performance in real-world applications.
Information
- Författare
- Karlin, Katarina, Amin, Diana
- Lärosäte / institution
- Stockholms universitet/Institutionen för data- och systemvetenskap
- Publiceringsdatum
- 2025
- Uppsatstyp
- Kandidat-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Yrkesexamen på avancerad nivå, Uppsala universitet/Avdelningen för systemteknik
Vigholm, Albin
Publicerad: 2026
Kandidat-uppsats, Jönköping University/Tekniska Högskolan
Rönnqvist, Emilia, Skoogh, Lovisa
Publicerad: 2026
Kandidat-uppsats, Högskolan i Skövde/Institutionen för informationsteknologi
Dargren, Calle
Publicerad: 2026
Yrkesexamen på avancerad nivå, Luleå tekniska universitet/Institutionen för ekonomi, teknik, konst och samhälle
Åström, Tuva, Nilsson, Matilda
Publicerad: 2026
Yrkesexamen på avancerad nivå, Luleå tekniska universitet/Institutionen för ekonomi, teknik, konst och samhälle
Nordlander, Jonas
Publicerad: 2026
M1-uppsats, Jönköping University/JTH, Avdelningen för datateknik och informatik
Seyhani Porshekoh, Artin
Publicerad: 2026