Uppsats

Lost in Translation, Found in Frames: Exploring LLMs for automatic translation of semantically annotated sentences

Kandidat-uppsats

Stockholms universitet/Avdelningen för datorlingvistik

Publicerad: 2026

Språk: Engelska

Sammanfattning

English FrameNet contains over 200,000 annotated sentences, while Swedish FrameNet is substantially smaller with around 9,000 sentences. This thesis investigates whether large language models can help bridge this resource gap by translating English FrameNet sentences into Swedish while preserving frame-semantic annotation. Two models, GPT-4.1 mini and Gemini 2.5 Flash, were given the task of translating 635 English sen- tences from three time-related frames (Change_event_time, Relative_- time, and Taking_time) into Swedish while preserving bracketed annota- tion labels and the target lexical unit. The outputs were evaluated using SacreBLEU, chrF, automatic target lexical unit matching, and a qual- itative error analysis inspired by the Multidimensional Quality Metrics framework. The results show that GPT-4.1 mini outperformed Gemini 2.5 Flash across all measures. GPT-4.1 mini achieved a SacreBLEU score of 76.63, a chrF score of 86.39 and a target preservation success rate of 63.0%, compared to Gemini 2.5 Flash’s 53.39, 77.18 and 38.6%, respec- tively. The qualitative analysis revealed different error profiles: GPT- 4.1 mini errors were evenly split between target and translation issues, while Gemini 2.5 Flash errors were dominated by incorrect target lexical units. These findings suggest that while LLMs can be useful for expanding frame-semantic resources, their performance varies significantly between models, and translation quality alone does not guarantee preservation of frame-semantic information.

Information

Lärosäte / institution
Stockholms universitet/Avdelningen för datorlingvistik
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.