Uppsats

Self-Supervised Fixed-Scene Adaptation for Object-Detection in Real-Time Surveillance: A Comparative Study of YOLO11 and RF-DETR

H

Chalmers tekniska högskola / Institutionen för elektroteknik

Publicerad: 2026

Språk: Engelska

Sammanfattning

This thesis investigates self-supervised fixed-scene adaptation for real-time objectdetectorsin an edge-computing surveillance context. While modern object-detectorsachieve strong results on general-purpose benchmarks, deployment in static camerascenes introduces distinct challenges: domain shift to a specific viewpoint, limitedavailability of scene-specific labels, and stringent compute and memory budgets ondevice.At the same time, the stationary background of surveillance footage providesexploitable structures, as do their temporal dependencies of video-frames. This studyconducts a comparative analysis of two state-of-the-art object detection architectures:the Transformer-dominant RF-DETR and the convolutional neural network (CNN)-dominant YOLO11. The thesis employs the 100Scenes dataset to represent a broadrange of surveillance environments. Experimental results demonstrate that RF-DETRconsistently achieves higher accuracy, smoother convergence, and greater robustnessthan YOLO11, albeit with higher hardware demands. In contrast, YOLO11 variants(with a frozen backbone) leverage the larger trainable capacity of the neck and headto enable high scene-specific adaptability. While this yields significant gains underquality labeling, it tends to increase sensitivity to imperfect pseudo-labels and therisk of overfitting. Furthermore, by systematically varying model scales, adaptationstrategies and environmental conditions the experimental design yields more than 3400distinct runs. First the work examines the extent to which smaller, specialized modelscan match the approach of substantially larger models. The experimental results showthat a small specialised model can compete with larger general models. Secondly, thestudy evaluated a proposed on-device self-supervised labeling strategy that integratesSAHI with a bidirectional implementation of ByteTrack. The proposed self-supervisedlabeling strategy provided reliable performance gains across all architectures andconfigurations, by recovering hard negatives, more specifically small, occluded andlow confidence instances. Thirdly, the study investigated background-context fusion(BF). It proved to be consistently improving the performance in general for RFDETR,while it proved inconsistent for YOLO11 and failed to increase robustnessagainst seasonality, suggesting it induced background-dependent overfitting. Finally,the study shows that all models being trained on a summer scene exhibit a decreasein relative performance compared with the non-adapted models during a seasonaldomain shift to a winter scene.

Information

Författare
Justad, Jacob
Lärosäte / institution
Chalmers tekniska högskola / Institutionen för elektroteknik
Publiceringsdatum
2026
Uppsatstyp
H
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.