Uppsats

Machine Failure Prediction in Industrial Production Using Machine Learning and Survival Analysis

Yrkesexamen på avancerad nivå

Umeå universitet/Institutionen för matematik och matematisk statistik

Publicerad: 2026

Språk: Engelska

Sammanfattning

This thesis investigates how machine learning and survival analysis can be used to support machine failure prediction in industrial production using operational data from one of Dometic’s production sites. The study uses historical maintenance records, power consumption data, and production volume data. The aim is to evaluate whether limited operational data, in the absence of detailed condition-monitoring sensor data, can provide useful predictive information for maintenance and production planning. Logistic Regression, Random Forest, XGBoost, and Bayesian Parametric Survival models were evaluated and compared. The classical machine learning models were formulated as binary classification models for predicting asset failure within 72 hours, while the survival analysis modeled time to failure using Weibull, Exponential, and Log-normal distributions. The final comparison focused on the practical usefulness of Logistic Regression and the Bayesian Survival model rather than treating them as directly comparable models. The models differed in terms of data structure, feature selection, and prediction formulation. The Bayesian Survival model achieved higher recall, precision, and F1-score at the fixed classification threshold, indicating that it identified a larger share of actual failures. However, this came at the cost of a very high false-positive rate, meaning that a large share of non-failure observations were incorrectly classified as failures. Logistic Regression was selected as the final model because it provided a more suitable balance between predictive performance, false-alarm control, interpretability, and practical usefulness. On the test set, the tuned Logistic Regression model achieved an accuracy of 69.30% and a ROC-AUC of 75.64%. Although the model missed a substantial share of actual failures, it produced fewer false alarms and showed a stronger ability to distinguish between failure and non-failure observations across classification thresholds. The most important predictive variables were related to recent and historical maintenance activity, together with selected power-consumption features. The results indicate that available operational data can support failure prediction, but that predictive performance is limited by the absence of detailed condition-monitoring data. The model should therefore be used as a decision-support tool rather than as an automatic trigger for maintenance decisions.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.