Uppsats

Phishing with Attention : A Study on AI-Enhanced Phishing Leveraging RAG

Yrkesexamen på avancerad nivå

Karlstads universitet/Institutionen för matematik och datavetenskap (from 2013)

Publicerad: 2025

Språk: Engelska

Sammanfattning

Recent advances in Artificial intelligence (AI), particularly in Natural Language Processing (NLP), following the breakthrough of Large Language Models (LLMs) capable of producing coherent, human-like text, are reshaping modern society. These capabilities offer unprecedented utility across various domains, from assistance in writing to software development. However, they also introduce critical cybersecurity challenges, such as highly deceptive phishing attacks that can exploit human vulnerabilities. One key concern is the growing accessibility and sophistication of LLMs, enabling adversaries to craft personalized spear-phishing messages with minimal effort. These novel threats can potentially evade traditional detection strategies and compromise user security at scale, underscoring the urgent need to examine and address the possible threats and potential deficiencies in existing detection systems. This thesis explores the offensive and defensive capabilities of LLMs by focusing on phishing email generation and detection. By collecting publicly available information via a developed Retrieval-Augmented Generation (RAG) pipeline, a generative LLM produced spear-phishing messages with and without the added contextual relevance. Subsequently, a fine-tuned LLM phishing classifier, trained solely on traditional datasets, was tested for its robustness against these AI-generated attacks. To accomplish this, a comparative quantitative experiment was conducted. The results revealed that the AI-generated phishing emails significantly challenged the classifier, which only correctly identified 13.27% of the new threats compared to the 99.54% of the traditional phishing messages. This stark contrast highlights the model's vulnerability— overfitting to traditional phishing patterns and signals underrepresented in the generated dataset— underscoring both the classifier's inadequacy for safeguarding users in real-world applications and the importance of incorporating synthetic AI samples into training. Ultimately, these findings emphasize the dual-use nature of LLMs and alarming shortcomings in existing phishing defenses. The study contributes to a deeper understanding of AI's evolving role and associated risks in cybersecurity and offers practical insights for improving phishing detection models and cybersecurity awareness programs.

Information

Författare
Martini, Rex
Lärosäte / institution
Karlstads universitet/Institutionen för matematik och datavetenskap (from 2013)
Publiceringsdatum
2025
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.