Sammanfattning

High-level software testing is often complex and highly situational, and there are areas where conventional methods of automation are simply not feasible. One such potential area is the post-mortem analysis of failed tests. This thesis explores the feasibility of using Large Language Models (LLMs) for automating the analysis of simulation-based test-scenarios within Tacsi, a tactical simulation environment developed by Saab AB. A system monickered ScenarioLens was implemented to summarise JSON-encoded recordings of these scenarios and to generate a conclusive briefing. Extensive experiments evaluated model size and prompting strategies, validated through a novel method of measuring accuracy tailored to the LLM-generated conclusions. Results demonstrate that LLMs are able to generalise well to this domain under certain constraints and offer promising avenues for integrating AI into high-level, domain-specific software testing workflows. Additionally, the thesis explores the use of multimodal LLMs for the same task, though this approach showed less potential compared to purely text-based methods.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.