Uppsats

Hide and Seek From 10 000 Feet: An empirical study of improvement methods for small object detection in aerial imagery under real-time constraints

Magister-uppsats

Linköpings universitet/Institutionen för teknik och naturvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

The increased usage of remote sensing imagery within surveillance, target tracking, and search operations has imposed a demand for object detection algorithms to find objects that are far from the sensor. This thesis explores a wide range of methods that aim to improve detection performance when objects appear small within images. By evaluating how different architectures, modifications, and preprocessing steps influence detection accuracy, inference time, and computational efficiency, questions regarding how state-of-the-art models perform, how improvements affect inference speed, and how targeted approaches influence the detectability of larger objects are addressed. The model architectures evaluated in this thesis include various iterations of the YOLO family (YOLOv5 to the recent YOLO26), detection transformers in the form of RF-DETR, as well as the modifications SO-YOLO and SO-DETR. All models were trained and evaluated on the AI-TOD dataset, with a supplementary evaluation on the VisDrone2019 dataset, providing metrics on small-sized objects through AI-TOD and objects of larger scale through VisDrone2019. In addition to the architectures, image slicing methods such as SAHI and SAFT, data augmentation methods, and model parameters were trialed in order to identify improvements. The findings of this study show that transformer-based approaches through RF-DETR yielded the best performance among baseline models, scoring an mAP@[0.5:0.95] value of 0.2 on the AI-TOD test set. Among the top-performing models, NMS-free designs were predominantly used. A substantial trade-off between accuracy and inference speed remains, as techniques such as upsampling, SAHI, and SO-YOLO significantly degrade runtime efficiency, rendering them impractical for real-time applications. Mosaic augmentation emerges as the most effective training strategy for improving general robustness; however, conventional augmentation methods primarily enhance localization stability rather than addressing the fundamental challenges of tiny object detection. Moreover, approaches targeted toward small objects often, but not always, degrade performance for larger objects.

Information

Lärosäte / institution
Linköpings universitet/Institutionen för teknik och naturvetenskap
Publiceringsdatum
2026
Uppsatstyp
Magister-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.