Uppsats

Deep learning model ensemble for remote sensing land use classification

Master-uppsats

Lunds universitet/Institutionen för naturgeografi och ekosystemvetenskap

Publicerad: 2025

Språk: Engelska

Sammanfattning

This study investigates deep learning approaches for automated land use classification from high-resolution remote sensing imagery, comparing Convolutional Neural Network (CNN) and Vision Transformer (ViT) architectures. Eight semantic segmentation models were evaluated on the HRSCD dataset containing 291 aerial image pairs (0.5m resolution) with four land use classes: building, agricultural, forest, and water. Models included CNN-based architectures (Deeplabv3+, HarDNet, STDC1-Seg50, STDC2-Seg50) and Transformer-based architectures (SETR, SegMenter, SegFormer-B1, SegFormer-B5). Transformer models demonstrated superior performance, with SegMenter achieving highest accuracy on Test Set I: Mean Pixel Accuracy (MPA) 94.22%, Mean Intersection over Union (MIoU) 78.67%, and Mean F1-Score 87.77%. All models showed degraded performance on Test Set II, with average decreases of 3.79%, 10.87%, and 9.68% for the three metrics respectively, indicating domain shift sensitivity. The optimal ensemble configuration combined three models (STDC2-Seg50, SegMenter, SegFormer-B1) using weighted averaging, achieving MPA 94.42%, MIoU 81.69%, and MF1-Score 89.71%. This represented improvements of 2.32%, 8.51%, and 5.81% over the best single model. Traditional classification methods (ISODATA, Maximum Likelihood, Random Forest, Support Vector Machine) achieved significantly lower comprehensive scores (31.20%-61.55%) compared to the ensemble's 88.61%. The superior performance of Transformers is attributed to self-attention mechanisms enabling better global context modeling. The ensemble approach successfully integrated complementary strengths: CNN models provided robust local feature extraction while Transformers contributed global semantic understanding. Water class classification remained challenging due to severe data imbalance. This research demonstrates that multi-model ensembles combining CNN and Transformer architectures provide a robust framework for operational land use mapping, though challenges persist in handling extreme class imbalance and cross-domain generalization.

Information

Författare
Yang, Jingfan
Lärosäte / institution
Lunds universitet/Institutionen för naturgeografi och ekosystemvetenskap
Publiceringsdatum
2025
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.