Uppsats
Enhanced Speech Emotion Recognition using MFCC Features and Convolutional Neural Networks
Kandidat-uppsats
Blekinge Tekniska Högskola/Institutionen för datavetenskap
Publicerad: 2024
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Background: Speech Emotion Recognition (SER) is a rapidly developing field in artificial intelligence that focuses on accurately interpreting human emotions through voice analysis. This thesis studies the improvement of SER systems using Mel-Frequency Cepstral Coefficients (MFCCs) for feature extraction and Convolutional Neural Networks (CNNs) for emotion classification. The major objective is to createa strong SER system that enhances the accuracy of detecting emotional subtleties in human speech. Objective: To achieve this objective, we gathered and preprocessed several emotional speech datasets, including CREMA-D, SAVEE, EMO-DB, TESS, and RAVDESS. These datasets were utilized to extract MFCC features, which were then input into CNN models specifically intended for emotion recognition. Our methodology included rigorous training and assessment methods, as well as strategies such as data standardization, augmentation, and model fine-tuning to improve performance. Methods: The experimental findings show that combining MFCCs with CNNsgreatly improves the accuracy of emotion identification systems. In both controlled and noisy conditions, our models performed admirably on a variety of criteria, including accuracy, recall, and F1 score. These findings demonstrate the resilience and practicality of the created SER system in real-world circumstances. Results: The experimental results demonstrate that combining MFCCs with CNNs significantly enhances the accuracy of emotion identification systems. The models performed well on various metrics, including accuracy, recall, and F1 score, under both controlled and noisy conditions. Specifically, the combined English and German-trained CNN model achieved the highest accuracy of 88%, surpassing the performance of models trained in a single language. Conclusion: This study establishes the efficacy of using MFCCs and CNNs for SER, highlighting the models’ ability to generalize across different datasets and conditions. These findings pave the way for future research, suggesting potential in exploring new machine learning techniques, diverse datasets, and multimodal emotion recognition methods.
Information
- Författare
- Gollapalli, Chandrasekhar, Raya, Nithin Chowdary
- Lärosäte / institution
- Blekinge Tekniska Högskola/Institutionen för datavetenskap
- Publiceringsdatum
- 2024
- Uppsatstyp
- Kandidat-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Shiva Olin, Harald
Publicerad: 2025
Kandidat-uppsats, Blekinge Tekniska Högskola/Institutionen för datavetenskap
Somula, Sai Deepak Reddy, Thadi, Giritanaya
Publicerad: 2025
Yrkesexamen på avancerad nivå, Blekinge Tekniska Högskola/Institutionen för datavetenskap
AL Bardan, Samer
Publicerad: 2025
Kandidat-uppsats, Lunds universitet/Matematisk statistik
Truong, Nancy
Publicerad: 2026
Kandidat-uppsats, KTH/Skolan för teknikvetenskap (SCI)
Johanson, Filip
Publicerad: 2026
Kandidat-uppsats, Karlstads universitet/Institutionen för hälsovetenskaper (from 2013)
Thoreson, Alice, Svensson, Björn
Publicerad: 2026