Uppsats
Scalable Vision–Language Machine Learning for Semantic Retrieval of Autonomous Driving Logs
H
Chalmers tekniska högskola / Institutionen för elektroteknik
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
This thesis studies scalable semantic retrieval of autonomous driving multi-viewvideos recorded with synchronized multi-camera systems. Using a subset of 3,502multi-view driving videos from NVIDIA’s PhysicalAI Autonomous Vehicles dataset,the work investigates text-to-multi-view video retrieval using natural-language queriesand learned cross-modal embeddings. Because the dataset does not contain pairedtextual descriptions, the proposed pipeline generates pseudo ground-truth captionsfrom sampled video frames using a pretrained vision-language model and extractsfrozen text and visual embeddings with jina-clip-v2. These generated captions providethe supervision used for training and evaluation without requiring manual annotation.Lightweight trainable alignment heads are then used to map text andvideo representations into a shared embedding space, while multi-view representationsare constructed through view-level and temporal aggregation. The resultsquantify the difference between single-view (front camera) and multi-view (frontand surrounding cameras) retrieval representations. Extending the representationfrom a single front-facing camera to six synchronized camera views increases Recall@5 from 71% to 85% and Recall@10 from 81% to 93%, indicating improvedseparation of ground-truth matches in the learned embedding space. In contrast,the LLM-based semantic similarity score changes only marginally, from 77 to 79,suggesting that both retrieval settings often retrieve semantically related drivingscenarios. The experiments further show that temporal sampling can be reducedconsiderably with only minor changes in retrieval performance, indicating substantialredundancy in densely sampled driving video. Since the supervision is derivedfrom automatically generated captions rather than human-annotated descriptions,the retrieval results should be interpreted with that limitation in mind. Overall,the thesis demonstrates that frozen pretrained encoders combined with lightweightfusion and alignment modules provide a computationally scalable approach for semanticretrieval of large-scale autonomous driving multi-view videos.
Information
- Författare
- Albarham, Mohammad, Berggren, Linus
- Lärosäte / institution
- Chalmers tekniska högskola / Institutionen för elektroteknik
- Publiceringsdatum
- 2026
- Uppsatstyp
- H
- Språk
- Engelska