Uppsats
Tool Use for Vision Language Models in Building Analysis of Street View Imagery
Master-uppsats
Linköpings universitet/Datorseende
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Pretrained vision-language models can be used for building analysis from street-view imagery, but many tasks still require fine-grained visual evidence that may be small, distorted, or partly occluded. This thesis investigates whether smaller open-weight vision-language models can improve building analysis by using external visual tools during inference. The study focuses on floor count estimation of buildings in Russian street-view panoramas. In each image, a red vertical line marks the target building, and the model must predict the number of visible above-ground floors. A tool-augmented VLM system is implemented in which the model can either answer directly or interact with a fixed set of visual tools, including zooming, segmentation, window detection, and marker drawing. The experiments compare direct prediction, prompted tool use, supervised fine-tuning on tool trajectories, and reinforcement learning as methods for learning when and how to use the tools. The results show that tool access alone is not sufficient for the tested smaller Qwen3-VL models. Prompted tool use decreases performance compared with direct prediction, indicating that the models struggle to control and interpret the tool interface without training. Supervised fine-tuning is the main source of improvement, making tool use more reliable by teaching valid tool calls and recurring interaction patterns. Warm-start reinforcement learning gives a smaller additional improvement for the best tool-enabled model, but does not outperform direct supervised fine-tuning without tools. In contrast, Gemini 3.1 Flash, used as a stronger proprietary reference model, benefits from tools even without task-specific training. Overall, the findings suggest that visual tools can support smaller VLMs in street-view building analysis, but only when tool use is learned as a reliable and selective interaction policy rather than added as a prompt-level capability.
Information
- Författare
- Pontén, Ture
- Lärosäte / institution
- Linköpings universitet/Datorseende
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
- Nyckelord
- ⌕agentic AI⌕Computer Vision⌕Reinforcement Learning⌕Gemini⌕vision–language models⌕multimodal AI⌕GRPO⌕OpenStreetMap⌕visual tool use⌕tool-augmented reasoning⌕street-view imagery⌕building analysis⌕floor count estimation⌕supervised fine-tuning⌕urban analytics⌕small vision-language models⌕Qwen3-VL⌕Russian street-view dataset⌕visual agents⌕open-weight models⌕Mapillary
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Suryanarayana Rao Prasanna, Navyashree
Publicerad: 2026
Master-uppsats, Linköpings universitet/Institutionen för systemteknik
Bülow, Gabriel
Publicerad: 2026
Master-uppsats, Uppsala universitet/Institutionen för informatik och media
Skeidsvoll Edén, Anniki
Publicerad: 2026
Master-uppsats, Malmö universitet/Institutionen för datavetenskap och medieteknik (DVMT)
Vrielink, Isabel
Publicerad: 2026
Master-uppsats, Lunds universitet/Industridesign
Nidadhalu Ramesh, Abishek
Publicerad: 2026
Master-uppsats, Blekinge Tekniska Högskola/Fakulteten för datavetenskaper
Bala, Neeraj
Publicerad: 2026