Uppsats
MemLogVR: Comparing Visual Guidance With and Without GenAI Voice Assistance in a VR Driving-and-Errands Navigation and Memory Task
Master-uppsats
Stockholms universitet/Institutionen för data- och systemvetenskap
Publicerad: 2026
Språk: Engelska
Sammanfattning
Introduction: Generative artificial intelligence assistants are increasingly available during everyday planning and task execution. They can hold task state supplied by surrounding software and repeat it back while a user is busy with another activity. This thesis reports a small comparative VR study of everyday incidental list recall during a driving-and-errands task. Research Question: The study asks how access to a state-aware push-to-talk voice assistant, added to shared visual current-target guidance, affects task completion, task performance, post-task recall, workload, user experience, and assistant-use patterns in a VR driving-and-errands task. Method: The study used a small non-randomized between-subjects design with deterministic balanced alternation. The baseline condition received a dashboard radar showing bearing and distance to the active target and a green ring at the active errand building. The assistant condition received the same visual guidance plus an OpenAI Realtime API push-to-talk voice assistant. The radar did not show the errand identity, whereas the assistant could verbally re-present the current errand during interaction. The intake workflow showed the scenario and errand list before the drive, but it did not announce a memory test or ask participants to memorize the list. Because the interface accepted only the current expected hotspot, the task removed the unaided initiation component and made post-task retrospective recall the primary memory outcome. Exploratory inferential tests and effect-size estimates are reported to make the comparison explicit, but they are interpreted cautiously because the sample is underpowered. Results: Sixteen participants started the task. Fourteen completed all ten errands and were included in completer outcome summaries (7 baseline, 7 assistant). Two assistant participants stopped after two of ten errands because of nausea, so completion among assigned participants was 7/7 in baseline and 7/9 in assistant. Among completers, item recall was higher in the assistant condition (median 4, range 2–6) than in baseline (median 2, range 0–5), while serial-position recall was the same by median (1 versus 1). The assistant condition also had a longer median completion time (660.7 s versus 593.4 s) and lower median mean speed (7.4 m/s versus 8.6 m/s). Exploratory tests did not provide a basis for confirmatory inference: for example, the item-recall comparison gave p = .103 with Cliff’s delta = .531, and completion among assigned participants gave Fisher’s exact p = .475. Discussion: Assistant logs showed 113 valid included turns, with the current errand named in 106 turns, one stale completed-errand mention, and no detected future-destination leakage. The recall difference is therefore best read as a small, sample-bound pattern observed under frequent current-item verbal representation, not as evidence about conversational AI in general. The platform, data pipeline, exploratory tests, and power estimates support a larger, counterbalanced, adequately powered follow-up rather than settling the comparison.
Information
- Författare
- Franke, Andreas
- Lärosäte / institution
- Stockholms universitet/Institutionen för data- och systemvetenskap
- Publiceringsdatum
- 2026
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska