Uppsats

Closing the Semantic Gap: Automated Test Suite Management using Multi-Agent Systems - A Design Science Research Study on Autonomous Test Suite Augmentation

Master-uppsats

Göteborgs universitet/Institutionen för data- och informationsteknik

Publicerad: 2026-06-29

Språk: Engelska

Sammanfattning

Requirement-test links can show that two artifacts are related, but they do notshow whether a test checks the behavior stated in a requirement. A linked test mayexercise the relevant feature while omitting the expected result or assertion needed toverify it. This thesis studies these cases as possible gaps in test evidence and presentsLanterns, a multi-agent research artifact for diagnosing such gaps and generatingcandidate tests.The study follows Design Science Research and evaluates Lanterns on EBT andLibEST datasets. The implementation separates traceability recovery from assertionlevel analysis because a related test is not necessarily a verifying test. The assertionreport records structured gap entries, which are passed to test generation togetherwith traceability and project evidence. This staged design links the diagnostic work inRQ1 with the gap-directed generation examined in RQ2, while keeping the evidenceand decisions from each stage available for review.Traceability predictions were compared with the dataset reference links, andassertion predictions were compared with a manually constructed ground truth thatwas checked through external review. Across two EBT runs, median recall was93.14% for traceability and 97.50% for assertion analysis, with median precision of57.93% and 54.87%, respectively. For LibEST, the median recall was 60.52% fortraceability and 68.81% for assertion analysis, while median precision was 40.00%and 35.63%. Most values were close between the two runs, although EBT assertionprecision showed greater variation. The results show that Lanterns recovered manypositive links, especially in EBT, but its false positives and missed links still requirehuman review.For RQ2, selected LibEST gap reports were used to generate candidate tests.Three cases were examined qualitatively to determine how closely each candidateaddressed its target gap. The assessment showed that gap-directed generation cantarget specific missing behaviors, although partial output and hallucinations remainrisks that require a technical review before adoption.The thesis contributes an inspectable workflow that connects traceability, verification analysis, and gap-directed test generation. The results support Lanterns asan aid for reviewing possible gaps in test evidence, rather than as a source of finalcoverage decisions or validated tests.

Information

Lärosäte / institution
Göteborgs universitet/Institutionen för data- och informationsteknik
Publiceringsdatum
2026-06-29
Uppsatstyp
Master-uppsats
Språk
Engelska