Uppsats

Exploring AI-driven Solutions for Automated Audio Descriptions of Videos : Region Stockholm

Kandidat-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2025

Språk: Engelska

Sammanfattning

This thesis explores the potential of using Artificial Intelligence (AI) to create audio descriptions for videos. AI has become more common in daily life and helps us in many areas. It supports tasks in education, public transport, healthcare and other sectors. In collaboration with Region Stockholm, the study explored the development of a system utilizing AI-driven solutions for automated audio descriptions of videos. The reason behind the project is that Region Stockholm currently struggles with producing audio described versions of videos. This has been met with complaints from Synskadades Riksförbund (SRF), and it is an issue that needs to be addressed due to legal requirements for accessibility which takes effect on June 28, 2025. The goal is to successfully develop, test and evaluate a prototype of the system, with the purpose of streamlining the creation process of audio descriptions. This should enhance accessibility for visually impaired individuals since it allows more videos to be audio described. This research faced several limitations and challenges, mainly related to budget constraints. To tackle these challenges, a prototype was implemented using free and low-cost AI services. The prototype includes two main features: description generation in SubRip Subtitle (SRT) format (subtitles) and speech generation using Text To Speech (TTS). The description generation is provided by Google’s Gemini 2.0 Flash, while the service for speech generation is provided by Azure OpenAI’s TTS model. A usability test of the prototype was conducted with participants representing the intended users at Region Stockholm. The results showed that participants were comfortable with certain features of the prototype and recognized its potential for future development. However, they also identified issues and areas that could be improved. For instance, the prototype could not sync the generated audio with the video. Furthermore, the prototype was demonstrated for the accessibility department of Sveriges Television (SVT), where they shared their feedback and considerations for further development. Overall, the insights gathered show that the prototype streamlined the process of creating audio descriptions by partial automation. This could lead to more videos to become audio described, thereby enhancing accessibility for visually impaired individuals.

Information

Lärosäte / institution
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publiceringsdatum
2025
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.