Uppsats

Beyond Standard Metrics: Cost-Sensitive Evaluation of Temporal Deep Learning for Predictive Maintenance

Master-uppsats

Stockholms universitet/Institutionen för data- och systemvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Class imbalance is a fundamental challenge in predictive maintenance, where failure events are rare but costly to miss. It remains unclear whether standard classification metrics reliably identify cost-optimal models when misclassification carries asymmetric financial consequences. This thesis investigates whether temporal deep learning architectures (LSTM and TCN) improve cost sensitive predictive performance compared to classical machine learning models on the SCANIA Component X dataset, and whether standard classification metrics reflect cost-based performance in imbalanced predictive maintenance settings. Five models were evaluated on over 23,000 vehicles: Logistic Regression,LSTM, TCN, SMOTE+LR, and SMOTE+XGBoost. Performance was assessed using both standard classification metrics (AUC-ROC, AUC-PR, recall, precision, F1) and cost-based evaluation via the expert-defined SCANIA cost matrix. To address the stochastic nature of neural network training, a multi-seed verification study was conducted across 30 independent initialisations per temporal model, with a principled quality filter applied prior to aggregation. TCN achieved the lowest mean test cost of 7.913 per vehicle, statistically significantly outperforming LSTM (8.086, p = 0.029) and surpassing published ensemble benchmarks. SMOTE augmentation improved upon baseline logistic regression (7.990 vs 8.748), but SMOTE+XGBoost failed to outperform the predict-all-zero baseline, suggesting that model complexity is not inherently beneficial when the training distribution is artificially altered. Temporal models consistently exhibited smaller validation-to-test cost gaps, indicating more transferable decision boundaries. Direct numerical comparison with published benchmarks should be interpreted with caution, as those studies employ the dataset’s native five-class labels whereas this thesis adopts binary classification. No consistent relationship was observed between any standard classification metric and cost per vehicle. A practitioner relying solely on AUC-ROC or AUC-PR would select a different model than cost-based evaluation identifies as optimal. In high-stakes industrial settings where misclassification costs are asymmetric, this discrepancy has direct financial consequences and suggests that cost-based evaluation should be treated as a primary objective rather than a downstream consequence of metric performance.

Information

Lärosäte / institution
Stockholms universitet/Institutionen för data- och systemvetenskap
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.