Sammanfattning

This thesis presents an empirical industrial comparison of single-agent and multi-agent Large Language Model (LLM)-based systems for supporting root cause analysis (RCA) of nightly test failures in an industrial test environment at Westermo. Nightly testing produces heterogeneous test-log data, and failures currently require manual inspection to identify relevant anomalies and possible root causes. This makes analysis time consuming, especially when failures involve several evidence sources or differ between test cases. LLM-based agents may support this work, but it remains unclear whether more complex agent architectures provide practical advantages over simpler alternatives in this context. To investigate this, two LLM-based RCA support systems were implemented and compared on the same task: one single-agent architecture and one multi-agent architecture. Both systems used tool-supported access to test metadata and logs to generate structured RCA reports for selected failure scenarios. The study was conducted as an exploratory industrial case study using a mixed-methods design. The evaluation included two test failure scenarios, survey responses and focus group feedback from six practitioners, and 120 system-level execution runs. The survey measured practitioner-rated accuracy, reasoning quality, fix realism, clarity, usefulness, and trust, while system measurements covered cost, duration, and consistency. The results show no consistent practitioner-perceived usefulness or perceived accuracy advantage for either architecture across the evaluated scenarios. However, the single-agent system required less time and lower cost to generate reports. This suggests that the single-agent architecture is the more practical baseline in the evaluated context, while the value of multi-agent architectures requires further study.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.