Uppsats
Cross-Modal Reasoning in Vision-Language Models for Robotics
Master-uppsats
Uppsala universitet/Institutionen för informationsteknologi
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Vision-Language-Action (VLA) models produce robot actions from visual observations and natural language instructions through an architecture that combines a vision encoder, a large language model (LLM), and an action head. During this process, the LLM develops intermediate representations that encode how the instruction relates to the visual scene. This thesis investigates whether these representations can be used to strengthen cross-modal interaction by enriching the visual features before action prediction. The investigation focuses on two openweight VLAs, OpenVLA and OpenVLA-OFT, which share the same pretrained backbone, and evaluates them on the LIBERO simulation benchmark and its perturbed variant, LIBERO-PRO. A diagnostic analysis using attention patterns, layer-wise masking, and wrong-instruction tests shows that both models are vision-dominated but differ in how strongly the instruction shapes the predicted action. The analysis also identifies layer regions in OpenVLA with distinct roles in visual and linguistic processing, motivating its selection as the base model. Based on these findings, three differentiable feedback mechanisms are proposed within a two-pass framework, each intervening at a different stage of the architecture. The first pass follows the base model's standard forward pass and caches signals from the layer regions identified in the diagnostic analysis. The second pass uses these signals to modulate only the visual representations before the base model predicts the final action. The base model is kept frozen, and only lightweight auxiliary modules are trained. The strongest results are obtained by averaging the predicted actions of two feedback variants that extract image-token hidden states from the LLM and feed them back to the patch embeddings of the vision encoders. This approach increases the average success rate across four LIBERO suites from 74.4% to 77.3%, with the gains mainly driven by more precise grasping. Evaluation under distribution shifts on LIBERO-PRO shows that the feedback strengthens execution when the task goal remains the same as during training, such as under rephrased instructions or changed object appearances, but cannot compensate when the goal changes or objects are relocated to new positions.
Information
- Författare
- Katranzopoulou, Maria
- Lärosäte / institution
- Uppsala universitet/Institutionen för informationsteknologi
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Kandidat-uppsats, Malmö universitet/Institutionen för datavetenskap och medieteknik (DVMT)
Andersson, Joel, Aldein Baradee, Noor
Publicerad: 2026
Magister-uppsats, Högskolan i Skövde/Institutionen för informationsteknologi
Dogan, Robin, Basunaid, Mazen
Publicerad: 2026
Kandidat-uppsats, Mittuniversitetet/Institutionen för data- och elektroteknik (2023-)
Muhammad, Al Jaber Al Shwali
Publicerad: 2024