Sammanfattning

This Bachelor's thesis evaluated how machine learning models trained in English perform in phishing detection for Scandinavian languages. Previous research has mostly focused on English-language datasets which makes it unclear how well these models generalize to other languages and whether machine translation can reduce the performance loss when the language changes. To evaluate this, a controlled experiment was conducted. Four models were tested: Naïve Bayes, Support Vector Machine (SVM), Random Forest, and Convolutional Neural Network (CNN). They were evaluated under four experimental conditions: an English baseline, zero-shot classification, roundtrip translation, and monolingual training. Due to limited public availability of phishing datasets in Swedish, Danish, and Norwegian, such datasets were created by machine translating an English phishing dataset. The results showed that all models performed well on English data, with an accuracy of up to around 99 percent. However, when applied directly to Scandinavian languages in the zero-shot condition, performance dropped drastically, falling to around 65 to 72 percent. Roundtrip translation improved the results compared to the zero-shot condition, but was unable to reproduce the original results observed in the baseline condition. When the models were trained and tested directly on Scandinavian languages in the monolingual condition, performance recovered to a level that was not as high as the baseline condition, but remained comparable. The study concluded that the language barrier constituted a considerable problem, as English trained models struggled to detect phishing in other languages. Language-adapted training data therefore appeared to be crucial for robust detection in cross-language scenarios. The findings indicated that machine translation-based approaches could only partially compensate for the performance degradation associated with language transfer.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.