Uppsats
Semantic Similarity of LLM Explanations in Software Engineering Tasks
Master-uppsats
Göteborgs universitet/Institutionen för data- och informationsteknik
Publicerad: 2026-06-29
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Large language models (LLMs) are increasingly used to generate natural-languageexplanations for software engineering (SE) tasks. However, it is still unclear howsimilar these explanations are across models when the task remains the same. Weaddress this gap in two SE contexts: Python code comprehension and explainingrequirements classifications. We use a controlled experimental design to compareGPT-4.1, Claude Sonnet 4, and Gemini 2.5 Flash under two prompt types: highlevel and detailed. In total, the experiment produces 240 explanations. We analyzethese explanations from three complementary perspectives: embedding-based semantic similarity, manual reasoning pattern analysis using a predefined codebook,and perceived similarity and quality through a human survey and an LLM-as-a-judgeevaluation. Our results show that explanations from different models are often closein meaning, but their similarity varies across model pairs, task types, and promptconditions. The reasoning pattern analysis shows that models also differ in how theystructure their explanations, and these differences depend on the task and prompttype. The quality evaluation of the selected explanation pairs shows that no modelis consistently better than the others across all dimensions, and that human andLLM-as-a-judge judgments do not always agree. Therefore, evaluating LLM explanations in SE should combine semantic similarity measures, reasoning-structureanalysis, and human or LLM-based quality judgments to support trustworthy useof these models in practice.
Information
- Författare
- Andersson, Simon, Wang, Qianyuan
- Lärosäte / institution
- Göteborgs universitet/Institutionen för data- och informationsteknik
- Publiceringsdatum
- 2026-06-29
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska