Uppsats

Computer Vision in Cooking Support Systems : Leveraging Multimodal LLMs for Dynamic, Open-Vocabulary, Real-Time Object Detection and Tracking

Kandidat-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2025

Språk: Engelska

Sammanfattning

A major limitation of current object detection models is their lack of dynamic identification capabilities, while large multimodal models (LMMs) often lack detailed spatial understanding and are too computationally demanding for real-time video applications. To bridge this gap, a prototype hybrid system, DOVRT-DT, is developed for Dynamic, Open-Vocabulary, Real-Time Object Detection and Tracking. The first research question explores the technical capabilities and limitations of DOVRT-DT. Evaluation was done on a video recorded in a kitchen environment. The results demonstrate strong dynamic identification and real-time capabilities but also highlight limitations related to inconsistencies in LMM formatting, object detection and tracking. As a potential implementation of this type of technology, the second research question explores how external factors might influence the business viability of a cooking support system within the context of an early adoption use case: supporting individuals with dementia in public care facilities in the EU. Public funding and public perception emerged as key uncertainties. Success will depend on how these factors evolve and whether manufacturers are willing to accept the associated risks.

Information

Lärosäte / institution
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publiceringsdatum
2025
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.