Uppsats
Anomaly Detection in Simulator Log Data Using Unsupervised Machine Learning Methods
Kandidat-uppsats
Linköpings universitet/Institutionen för datavetenskap
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Anomaly detection is an important task in industrial systems where large volumes of log data are generated continuously. In complex simulation environments, such as flight simulator systems, manual inspection of logs is time-consuming and often impractical. This creates a need for automated methods capable of identifying abnormal system behavior without relying on labeled data. This thesis investigates the effectiveness of two unsupervised machine learning methods, Isolation Forest and K-Means Clustering, for detecting anomalies in flight simulatorlog data. The study is based on approximately 3,300 unlabeled log files extracted from a simulation environment used for development and testing of complex systems. The logdata is transformed into numerical features capturing structural and behavioral characteristics such as event frequencies, error counts and log length. Both models are evaluated using a combination of anomaly scoring, distributional analysis and manual expert validation. In addition, a supervised validation experiment using a Support Vector Machine (SVM) with cross-validation is conducted to assess the consistency of the detected anomalies. The results show that both Isolation Forest and K-Means Clustering are capable of identifying meaningful anomalies in the dataset. Isolation Forest achieves higher precision and primarily detects globally rare and structurally deviating observations. K-Means Clustering identifies anomalies based on deviations from cluster structures and recurring system patterns but shows lower precision. The manual evaluation confirms that several detected anomalies correspond to meaningful system issues such as process interruptions and communication failures. The supervised validation showed that the anomalies were difficult to clearly separate from normal behavior, suggesting that the dataset contains varied and overlappinganomaly patterns. Overall, the findings suggest that Isolation Forest was more precise for identifying meaningful anomalies in this dataset, while both methods demonstrated the ability to detect abnormal system behavior in unlabeled industrial log data.
Information
- Författare
- Rydell, Ingrid
- Lärosäte / institution
- Linköpings universitet/Institutionen för datavetenskap
- Publiceringsdatum
- 2026
- Uppsatstyp
- Kandidat-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Kandidat-uppsats, KTH/Hälsoinformatik och logistik
Bergh Lagerqvist, Joel, Lund, Mattias
Publicerad: 2026
Kandidat-uppsats, KTH/Skolan för teknikvetenskap (SCI)
Persson, Elias, Råhlén, Gustav
Publicerad: 2026
Kandidat-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Selling, Felix
Publicerad: 2024
Master-uppsats, Linköpings universitet/Institutionen för datavetenskap
Hedström, Johannes
Publicerad: 2025
Kandidat-uppsats, KTH/Medicinteknik och hälsosystem
Iglesias, Sofia, Chinzorig, Anujin
Publicerad: 2026
Kandidat-uppsats, Mälardalens universitet/Institutionen för datavetenskap och datateknik
Hanebring, Daniel
Publicerad: 2026