Uppsats

IMPROVING POLLEN SPECIES CLASSIFICATION USING DEEP LEARNING AND SIZE FEATURES

Kandidat-uppsats

Lunds universitet/Matematisk statistik

Publicerad: 2026

Språk: Engelska

Sammanfattning

Automated pollen classification is challenging because visually similar species may differ only by subtle morphological characteristics. This thesis investigates deep-learning models for classifying microscopic pollen-grain images from ten plant species using RGB image crops together with two segmentation-derived size measurements: the major and minor axes of the pollen grain. The study builds on work by Jurdell (2025) and uses the same ResNet18-based image-and-size model in Python/PyTorch to establish a comparable baseline. This thesis evaluates the effects of missing-size handling, validation-based model selection, class-balancing augmentation, image-feature representation, focal loss, and backbone architecture. The reproduced baseline, denoted R512-Base, achieved 76.8% test accuracy and 76.2% macro recall. Species-wise mean imputation of missing size features and the change in image representation improved performance. The best overall model replaced Resnet with ConvNeXt-Tiny, achieved 85.2% test accuracy and 85.3% macro recall. The improvement was not uniform across species. The most persistent difficulty was the separation of Crepis capillaris and Hypochaeris radicata. In the baseline model, Crepis Capillaris had zero recall; the final ConvNeXt-Tiny model substantially increased the recall of Crepis capillaris to 38.5%. Thus, the main limitation was not poor performance across all species, but concentrated uncertainty around a small number of visually similar and flower-sensitive classes. A sensitivity analysis using 16 alternative flower-level splits for Crepis capillaris and Hypochaeris radicata showed that the final model remained strong on average, with mean accuracy 84.2% and mean macro recall 85.2%. However, the recalls of these two species varied substantially depending on which flowers were held out for testing, whereas the remaining species were comparatively stable. These findings support the use of flower-level evaluation for pollen classifiers and show that aggregate accuracy should be reported together with class-wise recall and error-structure analysis.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.