Uppsats

Anomaly Detection in Simulator Log Data Using Unsupervised Machine Learning Methods

Kandidat-uppsats

Linköpings universitet/Institutionen för datavetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Anomaly detection is an important task in industrial systems where large volumes of log data are generated continuously. In complex simulation environments, such as flight simulator systems, manual inspection of logs is time-consuming and often impractical. This creates a need for automated methods capable of identifying abnormal system behavior without relying on labeled data. This thesis investigates the effectiveness of two unsupervised machine learning methods, Isolation Forest and K-Means Clustering, for detecting anomalies in flight simulatorlog data. The study is based on approximately 3,300 unlabeled log files extracted from a simulation environment used for development and testing of complex systems. The logdata is transformed into numerical features capturing structural and behavioral characteristics such as event frequencies, error counts and log length. Both models are evaluated using a combination of anomaly scoring, distributional analysis and manual expert validation. In addition, a supervised validation experiment using a Support Vector Machine (SVM) with cross-validation is conducted to assess the consistency of the detected anomalies. The results show that both Isolation Forest and K-Means Clustering are capable of identifying meaningful anomalies in the dataset. Isolation Forest achieves higher precision and primarily detects globally rare and structurally deviating observations. K-Means Clustering identifies anomalies based on deviations from cluster structures and recurring system patterns but shows lower precision. The manual evaluation confirms that several detected anomalies correspond to meaningful system issues such as process interruptions and communication failures. The supervised validation showed that the anomalies were difficult to clearly separate from normal behavior, suggesting that the dataset contains varied and overlappinganomaly patterns. Overall, the findings suggest that Isolation Forest was more precise for identifying meaningful anomalies in this dataset, while both methods demonstrated the ability to detect abnormal system behavior in unlabeled industrial log data.

Information

Författare
Rydell, Ingrid
Lärosäte / institution
Linköpings universitet/Institutionen för datavetenskap
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.