Uppsats

Evaluating Channel Importance in Deep Learning-based Multimodal Cytopathological Image Classification

Master-uppsats

Uppsala universitet/Avdelningen Vi3

Publicerad: 2025

Språk: Engelska

Sammanfattning

Deep learning-based image classification has the potential to provide efficient AI-assisted cancer diagnostics. In particular, multimodal imaging enables neural networks to utilize complementary information that is not accessible through single-modality processing. The network architecture can fuse multimodal image data at different stages: early fusion combines all modalities at the input level, while late fusion merges the extracted features after separate modality-specific processing. This integration presents challenges for interpretability, particularly in tracing how individual image channels within each modality influence the network’s predictions. To address this, we propose a comprehensive analysis strategy that combines perturbation-based and gradient-based methods to quantify channel contributions. By evaluating the importance of individual image channels, this approach enables informed and targeted selection of channels and modalities, reduces redundancy, and holds potential for uncovering diagnostically relevant spectral features. We conduct the analysis on a 19-slide multimodal oral cancer dataset comprising digitized whole-slide cytology images acquired through brightfield and fluorescence microscopy. We aggregate and evaluate the significance of each of the seven input channels and systematically compare relevance patterns between early and late fusion strategies. Depending on the fusion strategy, we observe that multiple channel selection configurations using only four to six input channels outperform the full seven-channel baseline in overall classification performance. Late fusion facilitates a more balanced distribution of contributions across both imaging modalities, whereas early fusion is predominantly driven by features from brightfield imaging. Regardless of the fusion strategy, we identify the spectral region around 517 nm emission in the fluorescence channels as offering particularly strong diagnostic value for accurate oral cancer image classification. Our results show substantial variability in model performance across patient-level test folds, highlighting the need to expand the dataset to improve the generalizability of the identified channel selection strategies to new data. Despite this limitation, we consistently demonstrate that channel-optimized multimodal models outperform their unimodal counterparts, reinforcing the benefit of integrating diverse visual information representations in computational diagnostics.

Information

Författare
Derner, Nora
Lärosäte / institution
Uppsala universitet/Avdelningen Vi3
Publiceringsdatum
2025
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.