Uppsats

Detecting LLM-Based Web Honeypots Using Machine Learning

Kandidat-uppsats

Malmö universitet/Institutionen för datavetenskap och medieteknik (DVMT)

Publicerad: 2026

Språk: Engelska

Sammanfattning

One of the shortcomings of low and medium interaction honeypots is the static nature of their responses. To address this, solutions using Large Language Models (LLMs) have been proposed to equip them with dynamic response generation capabilities, serving to decrease detectability and consequently, improve attacker engagement. However, to ensure the efficacy of the systems, there is a need for rigorous evaluation of their detectability. While previous research has demonstrated machine learning classifiers capable of identifying static honeypots, this approach has yet to be applied to LLM-powered honeypots. Therefore, this thesis investigates the feasibility of utilizing machine learning technology to identify characteristics of LLM web honeypot responses for fingerprinting purposes. By applying novel evaluation techniques, we aim to provide developers and honeypot administrators with a tool for the assessment of honeypot detectability. To this end, a set of 1,597 URIs was collected and served as the basis for compiling a dataset. Legitimate responses, alongside generated responses from two open source honeypots (Galah, Beelzebub), were gathered and used to train a random forest classifier. The model proved capable of distinguishing between legitimate and honeypot responses, where response latency and lack of credible headers were shown to be of particular importance. We present a framework for fingerprinting LLM honeypots which, by delivering actionable results, could be used during both honeypot development and configuration.

Information

Lärosäte / institution
Malmö universitet/Institutionen för datavetenskap och medieteknik (DVMT)
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.