Uppsats
Exploring the use of Vision Language Models in End-to-End Motion Planning : Long-Tail Motion Planning for Autonomous Driving
Master-uppsats
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publicerad: 2025
Språk: Engelska
Sammanfattning
As autonomous driving systems are slowly becoming a reality, ensuring robust and reliable operation remains critical, particularly in safety-sensitive environments. Traditional motion planning approaches, which rely heavily on supervised learning from curated datasets, often struggle when faced with scenarios that deviate from their training distribution. This limits their ability to generalise to unexpected or rare driving conditions. Pre-trained foundation models, such as large language models and vision-language models, present a compelling alternative. Trained on vast and diverse datasets, these models offer powerful generalisation and reasoning capabilities that may precisely address the shortcomings of current data-driven motion planners. This thesis investigates the use of vision-language models for trajectory generation in both standard and more complex, challenging driving scenarios. We evaluate these models using both the nuScenes and NAVSIM benchmarks, focusing initially on their zero-shot capabilities and later on their performance after fine-tuning on task-specific data. However, due to the limited availability of representative edge-case scenarios and their inability to match the performance of State-of-The-Art planning systems, we propose an alternative role for these models. Rather than using them to directly generate motion plans, we explore their utility as post hoc evaluators, assessing the quality of expert-generated trajectories and identifying potential failure cases. Results show that while vision-language models demonstrate strong scene understanding and semantic reasoning, they frequently fail to produce trajectories that are spatially and temporally consistent or safe to execute. Challenges such as inconsistent output formatting, limited spatial grounding, and difficulties in interpreting multi-frame input highlight their limitations that currently hinder their deployment in motion planning. Nevertheless, this work suggests that foundation models hold significant promise as interpretable and modular components in autonomous driving pipelines. Their ability to provide contextual reasoning and semantic evaluation positions them well for complementary tasks such as trajectory validation, anomaly detection, and failure explanation, domains where transparency and generalisation are essential.
Information
- Författare
- Edmeades, Tobias
- Lärosäte / institution
- KTH/Skolan för elektroteknik och datavetenskap (EECS)
- Publiceringsdatum
- 2025
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Saleh, Abdelrahman
Publicerad: 2026
Yrkesexamen på avancerad nivå, Uppsala universitet/Avdelningen för systemteknik
Vigholm, Albin
Publicerad: 2026
Kandidat-uppsats, Högskolan i Gävle/Avdelningen för datavetenskap och samhällsbyggnad
Vambe, Vimbainaishe
Publicerad: 2026
Kandidat-uppsats, Högskolan i Halmstad/Akademin för informationsteknologi
Johansson, Nathalie, Jonsson, Liam
Publicerad: 2026
Kandidat-uppsats, Uppsala universitet/Institutionen för informatik och media
Olsson, Lukas, Arvidson, Jarl
Publicerad: 2026
Yrkesexamen på avancerad nivå, Uppsala universitet/Avdelningen för beräkningsvetenskap
Carlsson, Jesper
Publicerad: 2026