Uppsats
Interpretable deep-learning models for automatic ECG classification
Yrkesexamen på avancerad nivå
Uppsala universitet/Avdelningen för systemteknik
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Deep neural networks (DNNs) have demonstrated expert-level accuracy in automated 12-lead ECG diagnosis, yet their "black-box" nature remains a major barrier to clinical adoption. For medical professionals to safely rely on artificial intelligence, the underlying decision-making processes must be transparent. This thesis addresses this challenge by combining inherently interpretable data representations with post-hoc explanation techniques to develop accurate and transparent diagnostic models. To achieve structural interpretability, a β-variational autoencoder (β-VAE) was trained on the MIMIC-IV-ECG dataset. The model successfully compressed 12-lead median beats into a low-dimensional space consisting of 32 latent factors. Visual and statistical evaluations confirmed that these factors can capture distinct clinical ECG measurements, including ventricular rate, QT interval, and PR interval. To test the adaptability of the learned latent space, the MIMIC-trained β-VAE was fine-tuned on the PTB-XL training dataset to extract latent factors. Four different classification architectures (logistic regression, XGBoost, TabPFNv2.5, and TabICLv2) were then trained on these extracted training factors and evaluated on the unseen test data. The results indicate that this dimensionality reduction preserves the vast majority of relevant diagnostic information. When evaluating 23 diagnostic subclasses, the TabICLv2 model achieved a Macro-AUC of 0.927, performing on par with a 101-layer deep residual network trained directly on raw signals (AUC 0.929). Across all 71 distinct diagnostic classes, the model achieved a Macro-AUC of 0.906 compared to the residual network's 0.925. This remaining performance gap was partly driven by rhythm-related diagnoses, as the extraction of median beats naturally discards temporal inter-beat dynamics. Furthermore, applying SHAP analysis made it possible to track how the more flexible classification methods used these latent factors in a non-linear manner. When detecting Left Bundle Branch Block (LBBB), the evaluated models consistently prioritized the exact same latent variable (Factor 8) across both the MIMIC-IV-ECG dataset and the external PTB-XL dataset. This cross-dataset and cross-architecture stability indicates that the classifiers independently learned to rely on this same underlying morphological information. Overall, this thesis shows that pairing a VAE latent space with detailed SHAP analysis enables the construction of flexible classification models whose decisions can be traced back to visual ECG changes. The slight drop in predictive performance compared to unconstrained black-box networks is a trade-off. However, if this transparency helps doctors verify the AI's reasoning, this approach could be a step toward integrating deep-learning models into clinical practice.
Information
- Författare
- Wernersson, Olle
- Lärosäte / institution
- Uppsala universitet/Avdelningen för systemteknik
- Publiceringsdatum
- 2026
- Uppsatstyp
- Yrkesexamen på avancerad nivå
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Kandidat-uppsats, KTH/Skolan för teknikvetenskap (SCI)
Johanson, Filip
Publicerad: 2026
H, Chalmers tekniska högskola / Institutionen för elektroteknik
Tomasson, Moa, Westerkull, Saga
Publicerad: 2026
Yrkesexamen på avancerad nivå, Uppsala universitet/Datorteknik
Hedenström, Johan
Publicerad: 2025
Kandidat-uppsats, Lunds universitet/Matematisk statistik
Hitzemann, Max
Publicerad: 2026
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Du, Mingtong
Publicerad: 2026
Master-uppsats, Lunds universitet/Institutionen för elektro- och informationsteknik
Müller, Arvid, Flynn Rosenberg, Elias
Publicerad: 2026