Uppsats

Synthetic Data Augmentation for Intrusion Detection : Evaluating WGAN-GP for Class Imbalance and Novel Attack Detection

M1-uppsats

Blekinge Tekniska Högskola/Institutionen för datavetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Background. Society has become increasingly reliant on digital infrastructure, and with it an increasing importance is placed on cyber security. Intrusion detection systems (IDSs) play a crucial role in modern cyber security in mitigating the risk of cyber-attacks. However, the effectiveness of IDS models is constrained by the availability and quality of training data. Network flow datasets can be difficult to acquire due to their proprietary and sensitive nature, and those that are available frequently suffer from large class imbalances. This research aims to explore options for solving these problems through synthetic data generation. Objectives. This research investigates whether the Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP) can be used to generate high-fidelity synthetic data to alleviate class imbalances in an IDS context, and whether this augmentation can lead to improved detection inpreviously unseen and rare attack types. Methods. A WGAN-GP was trained per attack label on the CIC-IDS2017 dataset. Generated synthetic data had its fidelity rigorously tested using the Kolmogorov-Smirnov metric (KS), Pearson’s correlation coefficient, Kernel Density Estimation plots (KDE), and Train on Synthetic, Test on Real (TSTR) evaluation. Labels for which generated synthetic data was acceptable were included in an augmented classification pipeline, which was compared to a baseline pipeline using an ensemble of classifiers consisting of a Logistic Regression (LR) model, an XGBoost model, and a Multilayer Perceptron (MLP) model. These were evaluated on classification metrics, calibration plots, low resource experiment, and Area Under the Receiver Operating Characteristics Curve (ROC-AUC). A withheld label experiment was conducted at the same time in order to evaluate generalization and unseen attack detection. Results. Synthetic data generation succeeded for five out of eight labels, with varying fidelity results across labels. Downstream classification results produced negligible differences between augmented and baseline pipelines, largely attributed to problematic network flow data characteristics as well as dataset limitations. Synthetic data did not provide a noticeable improvement in previously unseen attacks. Conclusions. The synthetic data produced by WGAN-GP did not yield any meaningful improvements for either IDS classification performance or generalizability for unseen attacks. Results suggest that particularly problematic network data features such as large outliers, heavy skews, near-zero value clusters, and multimodality present significant challenges for the WGAN-GP generator and achieving high-fidelity synthetic data in this type of context likely requires extensive feature- and label-specific tuning. Future work should explore more targeted feature-engineering, more aggressive augmentation approaches, a hybrid WGAN-GP approach that combines global and class-specific generation, and evaluating datasets with less discriminative feature separation.

Information

Författare
Ringström, John
Lärosäte / institution
Blekinge Tekniska Högskola/Institutionen för datavetenskap
Publiceringsdatum
2026
Uppsatstyp
M1-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.