Uppsats
LLM-based Log Analysis - Evaluating LLM-Based Log Analysis in Industrial CI Environ ments: Technical Metrics, Practitioner Perspectives and Per ceived Workflow Impact
Master-uppsats
Göteborgs universitet/Institutionen för data- och informationsteknik
Publicerad: 2026-06-29
Språk: Engelska
Sammanfattning
Continuous Integration (CI) pipelines generate large volumes of complex logs thatdevelopers must analyze to diagnose failures and identify root causes. As software systems become increasingly distributed, manual log analysis has become timeconsuming, cognitively demanding, and highly dependent on practitioners’ expertise.Recent advances in Large Language Models (LLMs) have created new opportunitiesfor automated log analysis and debugging support. However, there remains a limitedunderstanding of how LLM-based log analysis tools should be evaluated and howpractitioners perceive their usefulness and influence on industrial CI environments.This thesis addresses that gap through an industrial case study conducted in collaboration with Bosch. The study follows a mixed-methods research design thatcombines qualitative and quantitative approaches, including a systematic literaturereview, technical evaluation, descriptive statistical analysis, and correlation analysesusing Spearman’s rank correlation and Kendall’s tau.The findings show that conventional text-overlap metrics, such as BLEU and ROUGE,are not suitable for evaluating log analysis tool outputs because they require a reference and cannot handle the reasoning-intensive nature of CI log analysis. LLM-as-ajudge metrics, specifically G-Eval and GPTScore, were identified as more appropriate because they support reference-free and reasoning-aware evaluation. However,the comparison between automated and practitioner evaluations revealed only weakalignment and limited discriminative power in automated scoring, indicating thatautomated evaluation cannot reliably replace human judgment. Practitioners werefound to evaluate outputs based on precise root-cause identification and whether theoutputs support debugging. The study further shows that practitioners perceive theLLM-based tool as useful for supporting CI troubleshooting by accelerating failureinvestigation, reducing manual log inspection, and providing interactive support,particularly for complex and unfamiliar failures. It also highlights risks related tohallucinations, over-reliance, and reduced critical reasoning.Finally, practitioners perceived the integration of an MCP-based agent into thedevelopment environment as useful for supporting workflow continuity by reducing context switching and enabling iterative conversational debugging. However,they viewed it as complementary to, rather than a replacement for, the standaloneweb-based tool. Together, the findings contribute to evaluation methodologies andempirical insights into how LLM-based log analysis tools can be assessed and considered for adoption in industrial CI workflows.
Information
- Författare
- Gupta, Pallavi, Hailu, Tesfu Tekleab
- Lärosäte / institution
- Göteborgs universitet/Institutionen för data- och informationsteknik
- Publiceringsdatum
- 2026-06-29
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
H, Chalmers tekniska högskola / Institutionen för data och informationsteknik
Tesfu Tekleab, Hailu, Gupta, Pallavi
Publicerad: 2026