Uppsats

Mowing with Meaning: Semantic Obstacle Detection for Autonomous Lawn Mowers Using VLMs

Master-uppsats

Stockholms universitet/Institutionen för data- och systemvetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Autonomous service robots operating in unstructured outdoor environments depend largely on reactive sensing and fixed-vocabulary perception models that can identify what is present in a scene but cannot reason about what is happening in it. Knowing that a region is classified as a person or an animal does not, by itself, determine the appropriate navigation action: the right response depends on the object's state, activity, and position, none of which a standard segmentation model can encode. The research question addressed in this thesis is whether a modular Vision-Language Model (VLM) based split-compute architecture is a suitable approach for semantic obstacle detection in this context. The proposed system deploys lightweight image segmentation on a resource constrained onboard computer and offloads visual reasoning to a nearby edge node running a VLM, which produces a structured natural-language response including a scene description, obstacle classification, and navigation action. This architecture was selected because it enables open-vocabulary visual reasoning on hardware that cannot support a full VLM locally. Suitability is assessed along two dimensions: the capability of the VLM layer to reason about obstacles in situations where semantic understanding beyond pixel-level classification is required, and whether its natural-language output improves the diagnosability of navigation errors compared to classification only systems. The system was deployed on a Husqvarna Automower robotic lawnmower and assessed through a multi-part evaluation on real robot footage and synthetic imagery. The pipeline correctly handles stationary obstacles and extends well beyond what segmentation alone can produce: it treats a running animal as a transient obstacle, recognises a picnic as a group activity and returns a contextually appropriate navigation decision, and differentiates ornamental from wild vegetation based on visual reasoning rather than a predefined class list. Dual-frame inference improved performance on moving-subject scenarios. The natural-language output made navigation errors directly attributable to specific model behaviours or prompt design choices, a diagnostic capability unavailable in segmentation-only systems. The results demonstrate the viability of the approach as a research prototype. The architecture matches or exceeds the segmentation baseline across all tested conditions and provides capabilities that are architecturally out of reach for fixed-vocabulary models. The natural-language output additionally improved diagnosability: errors were attributable to identifiable model behaviours or prompt design choices rather than remaining opaque. The remaining open questions around prompt portability and temporal reasoning are defined as the main directions for future work.

Information

Författare
Arbey, Louis
Lärosäte / institution
Stockholms universitet/Institutionen för data- och systemvetenskap
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.