Uppsats

Machine Learning-Based Detection of MCP Attacks

Yrkesexamen på avancerad nivå

Blekinge Tekniska Högskola/Institutionen för datavetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Background. The Model Context Protocol (MCP) is a new and emerging technology that extends the functionality of large language models, thereby improving workflows, but it also exposes users to a new attack surface. Several studies have highlighted related security flaws, but MCP attack detection remains underexplored. Objectives. To address this research gap, this study aims to develop and evaluate a range of supervised machine learning approaches for detecting malicious MCP tool descriptions. The study considers three scenarios: (1) a binary classification task distinguishing malicious from benign tools, (2) a multiclass classification task identifying the attack type while separating benign from malicious tools, and (3) a multiclass classification task with redundant classes combined, resulting in fewer classes. Methods. This research applied the Design Science Methodology. During the initial problem-identification phase, a literature review was conducted to identify the problem and its underlying causes. The second phase began with a more in-depth literature review of existing solutions and state-of-the-art research, and the primary goal of this phase was to define the artifact. In the third phase, the artifact was evaluated to determine whether the research questions had been properly addressed. For evaluation, both traditional and deep-learning models were applied to the tasks, and a rule-based approach was used as a baseline for comparison. Results. The results indicate that several of the developed models achieved 100% F1-score on the binary classification task. In the multiclass scenario, the BERT and SVC models performed best, with mean F1-scores of 81.9% and 81.0%, respectively. Confusion matrices were also used to visualize the full distribution of predictions often missed by traditional metrics, providing additional insight for selecting the best-fitting solution in real-world scenarios. Conclusions. This study presents an addition to the MCP defense area, showing that machine learning models can perform exceptionally well in separating malicious and benign data points. Furthermore, the study shows that these models can outperform traditional rule-based solutions that currently exist in the field.

Information

Lärosäte / institution
Blekinge Tekniska Högskola/Institutionen för datavetenskap
Publiceringsdatum
2026
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.