Uppsats

Bootstrapping av kunskapsinhämtning med hjälp av en fröontologi och små språkmodeller

Master-uppsats

Jönköping University/Tekniska Högskolan

Publicerad: 2026

Språk: Svenska

Sammanfattning

Low-resource industrial domains face several challenges in knowledge acquisition, including limited annotated data, complex domain knowledge expressions, and the high deployment cost of Large Language Models (LLMs). To address these problems, this thesis studies a knowledge acquisition method that combines a seed ontology, data augmentation, and fine-tuned Small Language Models (SLMs), and applies these to the casting domain. This study mainly focuses on two information extraction tasks: Named Entity Recognition (NER) and Relation Extraction (RE). To expand the limited annotated data, this thesis designs two data augmentation methods: contextual rewriting and coreference aware repair, and ontology guided text augmentation. This study uses the augmented data to fine-tune five instruction-tuned SLMs for NER and RE, and compares their inference results with non-fine-tuned SLMs, a fine-tuned Large Language Model (LLM), and a non-fine-tuned LLM. The results show that both fine-tuning and data augmentation significantly improve the performance of SLMs in the casting domain. On the fully augmented dataset, Llama-3.2-3B achieves the best overall performance among the five fine-tuned SLMs in the NER stage. Its F1 score increases from 0.351 before fine-tuning to 0.786 after fine-tuning, an increase of 43.5 percentage points, and is only slightly lower than the fine-tuned GPT-4.1 score of 0.795. In the RE stage, Llama-3.2-3B also achieves the highest SLM F1 score, increasing from 0.306 to 0.940, an improvement of 63.4 percentage points, and outperforming the fine-tuned GPT-4.1 score of 0.831. The data augmentation experiments further show that combining the two augmentation methods gives the best results, improving NER F1 from 0.426 to 0.707 and RE F1 from 0.614 to 0.927. When used separately, contextual rewriting and coreference-aware repair perform better on NER, while the two methods show similar effects on RE. In addition, this thesis tests the complete NER-to-RE pipeline on new texts that were not used during training. The results show that fine-tuned small language models can support practical domain knowledge extraction. However, expert review also shows that some extracted triples are still semantically unreasonable and require further domain validation. Overall, this thesis shows that seed ontology and data augmentation can effectively support knowledge acquisition in low-resource industrial domains. When combined with fine-tuned SLMs, they can improve domain-specific NER and RE performance while reducing dependence on large external language models.

Information

Lärosäte / institution
Jönköping University/Tekniska Högskolan
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Svenska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.