Sammanfattning

As computer vision becomes increasingly integrated into our society the need to acquire large annotated datasets increases. Synthetic datasets provide a cheaper method of automatically generating annotated data but introduce problems such as domain shift. Furthermore, synthetic dataset are typically domain-specific limiting their reusability across different tasks. This thesis investigates the impact of non-domain-specific synthetic data for training deep learning models in the task of automatic logo detection. Synthetic data was generated using the Unity game engine together with the Unity Perception Package to create varied, annotated images from three general environments: a city, a forest and a room. Using the generated datasets and the real-world Logos in the Wild (LITW) dataset, nine computer vision models were trained utilizing the YOLO11n architecture. These models were evaluated on a LITW and a synthetic dataset and benchmarked against each other. The synthetic models were then fine-tuned with the LITW dataset to explore any further impact on performance. The results showed that models trained solely on non-domain-specific synthetic data performed poorly when benchmarked against the real-world data. However, the fine-tuned models displayed competitive performance, in some cases outperforming the LITW model’s performance especially when trained on a more varied synthetic dataset. These findings suggest that while non-domain-specific synthetic data may not be sufficient in by itself, it can serve as a valuable source in fine-tuning and reduce the need for large amounts of annotated real data.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.