Uppsats
Bootstrapping av kunskapsinhämtning med hjälp av en fröontologi och små språkmodeller
Master-uppsats
Jönköping University/Tekniska Högskolan
Publicerad: 2026
Språk: Svenska
Sammanfattning
Low-resource industrial domains face several challenges in knowledge acquisition, including limited annotated data, complex domain knowledge expressions, and the high deployment cost of Large Language Models (LLMs). To address these problems, this thesis studies a knowledge acquisition method that combines a seed ontology, data augmentation, and fine-tuned Small Language Models (SLMs), and applies these to the casting domain. This study mainly focuses on two information extraction tasks: Named Entity Recognition (NER) and Relation Extraction (RE). To expand the limited annotated data, this thesis designs two data augmentation methods: contextual rewriting and coreference aware repair, and ontology guided text augmentation. This study uses the augmented data to fine-tune five instruction-tuned SLMs for NER and RE, and compares their inference results with non-fine-tuned SLMs, a fine-tuned Large Language Model (LLM), and a non-fine-tuned LLM. The results show that both fine-tuning and data augmentation significantly improve the performance of SLMs in the casting domain. On the fully augmented dataset, Llama-3.2-3B achieves the best overall performance among the five fine-tuned SLMs in the NER stage. Its F1 score increases from 0.351 before fine-tuning to 0.786 after fine-tuning, an increase of 43.5 percentage points, and is only slightly lower than the fine-tuned GPT-4.1 score of 0.795. In the RE stage, Llama-3.2-3B also achieves the highest SLM F1 score, increasing from 0.306 to 0.940, an improvement of 63.4 percentage points, and outperforming the fine-tuned GPT-4.1 score of 0.831. The data augmentation experiments further show that combining the two augmentation methods gives the best results, improving NER F1 from 0.426 to 0.707 and RE F1 from 0.614 to 0.927. When used separately, contextual rewriting and coreference-aware repair perform better on NER, while the two methods show similar effects on RE. In addition, this thesis tests the complete NER-to-RE pipeline on new texts that were not used during training. The results show that fine-tuned small language models can support practical domain knowledge extraction. However, expert review also shows that some extracted triples are still semantically unreasonable and require further domain validation. Overall, this thesis shows that seed ontology and data augmentation can effectively support knowledge acquisition in low-resource industrial domains. When combined with fine-tuned SLMs, they can improve domain-specific NER and RE performance while reducing dependence on large external language models.
Information
- Författare
- Yang, Yi, Yuan, Tingjun
- Lärosäte / institution
- Jönköping University/Tekniska Högskolan
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Svenska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Luleå tekniska universitet/Drift, underhåll och akustik
Roohani, Muhammad Ammar
Publicerad: 2024
Kandidat-uppsats, Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Bergman, David, Saleh, Moayad
Publicerad: 2026
Magister-uppsats, Linköpings universitet/Institutionen för teknik och naturvetenskap
Bergström, Elin, Svensson, Tobias
Publicerad: 2026
H, Chalmers tekniska högskola / Institutionen för data och informationsteknik
HANI, SALAM, LINDER, JONATHAN
Publicerad: 2025