Uppsats
Exploring AI-driven Solutions for Automated Audio Descriptions of Videos : Region Stockholm
Kandidat-uppsats
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publicerad: 2025
Språk: Engelska
Sammanfattning
This thesis explores the potential of using Artificial Intelligence (AI) to create audio descriptions for videos. AI has become more common in daily life and helps us in many areas. It supports tasks in education, public transport, healthcare and other sectors. In collaboration with Region Stockholm, the study explored the development of a system utilizing AI-driven solutions for automated audio descriptions of videos. The reason behind the project is that Region Stockholm currently struggles with producing audio described versions of videos. This has been met with complaints from Synskadades Riksförbund (SRF), and it is an issue that needs to be addressed due to legal requirements for accessibility which takes effect on June 28, 2025. The goal is to successfully develop, test and evaluate a prototype of the system, with the purpose of streamlining the creation process of audio descriptions. This should enhance accessibility for visually impaired individuals since it allows more videos to be audio described. This research faced several limitations and challenges, mainly related to budget constraints. To tackle these challenges, a prototype was implemented using free and low-cost AI services. The prototype includes two main features: description generation in SubRip Subtitle (SRT) format (subtitles) and speech generation using Text To Speech (TTS). The description generation is provided by Google’s Gemini 2.0 Flash, while the service for speech generation is provided by Azure OpenAI’s TTS model. A usability test of the prototype was conducted with participants representing the intended users at Region Stockholm. The results showed that participants were comfortable with certain features of the prototype and recognized its potential for future development. However, they also identified issues and areas that could be improved. For instance, the prototype could not sync the generated audio with the video. Furthermore, the prototype was demonstrated for the accessibility department of Sveriges Television (SVT), where they shared their feedback and considerations for further development. Overall, the insights gathered show that the prototype streamlined the process of creating audio descriptions by partial automation. This could lead to more videos to become audio described, thereby enhancing accessibility for visually impaired individuals.
Information
- Författare
- Wang, Ziang, Osman, Haron
- Lärosäte / institution
- KTH/Skolan för elektroteknik och datavetenskap (EECS)
- Publiceringsdatum
- 2025
- Uppsatstyp
- Kandidat-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Kandidat-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Löfgren, Nils
Publicerad: 2026
Kandidat-uppsats, Högskolan i Halmstad/Akademin för informationsteknologi
Fawal, Raghad
Publicerad: 2026
Kandidat-uppsats, Mälardalens universitet/Akademin för ekonomi, samhälle och teknik
Sauleskalne, Patricija, Tigerbacke, Fideli
Publicerad: 2026
Kandidat-uppsats, Karlstads universitet/Handelshögskolan (from 2013)
Alvenborg, Tilda
Publicerad: 2026
Kandidat-uppsats, KTH/Hälsoinformatik och logistik
Abdulnoor, Tia, Bygde, Thea
Publicerad: 2026
Kandidat-uppsats, Högskolan i Skövde/Institutionen för handel och företagande
Kling, Ellen, Rakh, Shilan
Publicerad: 2026