Uppsats

LLM-assisted SOC Alert Analysis : A Design Space Evaluation of system variants

Master-uppsats

Linköpings universitet/Institutionen för datavetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Analysts in Security Operations Centers (SOCs) face large volumes of security alerts, manyof which require repetitive contextual investigation before they can be classified as malicious or benign. Large language model (LLM) agents offer a possible way to automateparts of this triage work, but it is unclear which contextual factors such agents need andhow these factors should be exposed to avoid misleading conclusions. This thesis investigates how an LLM-based agent can support selected aspects of SOC alert analysis bycombining alert data, surrounding evidence, and structured investigation guidance.The study designs and evaluates a SOC alert-analysis pipeline in a controlled SOC-likeenvironment. The pipeline is based on contextual factors from prior work on how SOC analysts investigate alerts, especially the factors described by Kersten et al. These factorswere implemented through Model Context Protocol (MCP) tools, forming the Kerstenfactor MCP (K-MCP) configuration, and refined through prompt changes, tool-descriptionchanges, and progressive disclosure of evidence. The evaluation uses a labelled alert setthat combines benign alerts generated in the lab environment with malicious alerts takenfrom the public MORDOR 2019 dataset. Verdict correctness is evaluated with F1, precision,and recall, while report quality is evaluated with MESSALA, an LLM-based report-qualityevaluation framework.In this controlled evaluation, the results show that the way contextual evidence is given tothe agent has a clear effect on performance. A final classification performance was summarized using three repeated runs of the selected scoped K-MCP configuration. This configuration, which excluded alert history and surrounding-alert context to reduce contextualoverreach, averaged F1 = 0.942 ± 0.028 on the 76-alert evaluation set. The base referenceresult achieved F1 = 0.711 ± 0.016. The improvement mainly came from reducing falsepositives on benign alerts: the base result produced 12 false positives, while the selectedscoped K-MCP configuration produced an average of 1.0 false positives. Both configurations detected 16 of 17 malicious alerts. MESSALA complemented the verdict metrics byshowing how the contextual variants affected the quality and grounding of the generatedreports.These findings suggest that, in a controlled SOC-like evaluation setting, LLM agents can bestudied as evidence-guided decision-support systems rather than as standalone classifiers.The contribution of this thesis is not to show that the system is ready for operational use.Instead, the thesis studies how selected contextual factors and tool-design choices affect anLLM-assisted alert-analysis pipeline. The results also show that context must be scopedcarefully: adding more contextual information is not always beneficial and can increasemisleading reasoning

Information

Lärosäte / institution
Linköpings universitet/Institutionen för datavetenskap
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska