Uppsats

Self supervised anomaly detection pipeline

Yrkesexamen på avancerad nivå

Blekinge Tekniska Högskola/Institutionen för datavetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Background. Anomaly detection in an industrial environment is highly relevant as a failure to detect anomalies could worsen a company's reputation and cost a great deal. Manual inspection takes time and is prone to human error. Machine learning based methods can prove an effective anomaly-detection aid either replacing manual detection or supplementing it to make it more effective. Visually based systems have demonstrated potential as a solution with low material costs with significant adaptability. Objectives. The objective of this thesis is to develop a method to detect anomalies, effective both in terms of manual effort required and inference cost. The aim is to minimize manual annotation and avoid the need for manual gathering of anomalous data. Methods. The proposed solution in this thesis is a self-supervised architecture where the first System, System-A, is based on DINOv2 and Softpatch which are trained on normal data. The second system, System-B is introduced, where similar anomalies are grouped together for faster damage type classification. The thesis evaluated the clustering models, K-means, DBSCAN, HDBSCAN, Affinity Propagation and a deep embedded clustering model, with semi-supervised and unsupervised training.To mitigate the high inference cost of automated anomaly detection, System-C is proposed. System-C aims to distill the patterns found by the heavy System-A (teacher) into a lighter GAN model (student), by utilizing the heatmaps created by System-A. Results. The results indicate a heavier teacher model can train a lighter student model, that generates heatmaps with visually high reconstruction quality. The adaptability of the GAN model is demonstrated as the distillation process results in a significantly smaller model with six times the speedup. Furthermore, this research highlights the potential for weaknesses in System-A being transferred to System-C. Finally the clustering showed the embedding space of each model where the deep embedded clustering showed more compact clusters while the Affinity Propagation and K-means had the highest classification scores. Conclusions. In conclusion, distilling a heavier model into a lighter one through the created heatmaps does reduce the inference cost and can create visually high quality heatmaps. However it's challenging to draw a definitive conclusion regarding the viability of model distillation with limited data. The clustering models showed potential for good clustering, where K-means, Affinity Propagation and deep embedded clustering were the ones that stood out. However it was inconclusive which of these models were the best due to their average performance and differentiation between clustering quality and classification score.

Information

Lärosäte / institution
Blekinge Tekniska Högskola/Institutionen för datavetenskap
Publiceringsdatum
2026
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.