Uppsats

Testing the Validity of Music Emotion Maps in Arousal-Valence Space through Music Information Retrieval

Kandidat-uppsats

Stockholms universitet/Institutionen för data- och systemvetenskap

Publicerad: 2025

Språk: Engelska

Sammanfattning

Music has the unique ability to transcend cultural and linguistic boundaries, forging emotional connections across diverse contexts. Through the use of machine learning techniques, Music Information Retrieval (MIR) algorithms enable the extraction of audio features such as tempo, instrumentalness, valence, and arousal. These audio features have diverse applications, including personalized music recommendations on streaming platforms and procedural content generation in video games. However, a significant challenge in Music Information Retrieval (MIR) and Music Emotion Recognition (MER) lies in the systems and research gap in regard to the accuracy of translating subjective emotional qualities, such as valence (positivity) and arousal (energy), from subjective human perceptions to objective computational models based on audio features. This gap is particularly relevant because these systems are widely employed in the music industry, especially in consumer-facing subscription models, where they influence listener habits and artist livelihoods. Given their broad implications, it is critical to assess the accuracy and validity of these models to determine their reliability. This study addresses this gap by examining the translation of subjective emotional attributes between a computational description model (Spotify’s audio analysis API) and a psychological description model (J. Russell’s valencearousal circumplex). By analyzing the relationship between Spotify’s audio features and human emotional responses, the research investigates how effectively computational models can capture and predict subjective emotional experiences. This is realized through the following research question: “Which combination of audio features, according to J. Russell’s valence-arousal model, most accurately determines the arousal and valence dimensions of a song using Spotify’s audio analysis API?” By addressing this question, the study bridges a critical gap in MIR/MER research by offering a systematic approach to better understanding the relationship between computational audio features and human emotional perception, paving the way for more accurate and reliable emotional prediction systems. This study employs a survey as its research strategy, using online questionnaires to collect quantitative data on emotional responses to Spotify tracks to assess the alignment between human emotions and Spotify’s audio features, providing a foundation for regression analysis to evaluate the significance of audio features in predicting emotional dimensions of valence and arousal, entailing in essence positivity and energy respectively. The results reveal that the audio features danceability and loudness are reliable predictors of valence, while acousticness and danceability consistently predict arousal. Loudness (R = .568, β = -.420) and danceability (R = .282, β = .567) showed significant correlations with valence, indicating that more danceable tracks tend to have a more positive emotional tone, while louder tracks are associated with a slightly more negative emotional valence. For arousal, acousticness (R = .618, β = -.508) and danceability (R = .303, β = .559) were the most consistent predictors. Acousticness demonstrated a strong inverse relationship with arousal, suggesting more acoustic tracks correlate with lower arousal levels. The findings suggested that Spotify’s valence and arousal values aligned well with participants’ emotional responses. Valence (R = .566, Coefficient = .512) and Arousal (R = .537, Coefficient = .705) demonstrated strong predictive power within J. Russell’s Valence-Arousal circumplex model. However, the relationship between audio features and emotional responses is not always linear, and some features such as instrumentalness and key yielded insignificant results. Therefore, while the algorithm shows promise in estimating emotional responses, it needs further refinement. Ethical concerns arise around the commercial use of MIR algorithms, particularly their potential to manipulate listener behavior and introduce bias, affecting artist discoverability and listener experience. While the study provides valuable insights into the potential of Music Information Retrieval (MIR) in estimating emotional responses, limitations such as the reliance on proprietary algorithms and a limited datasets highlight the need for future research at a larger scale to refine these findings and further advance the field of MIR. Looking ahead, future studies could explore advanced machine learning techniques including neural networks and natural language processing for lyrical content to provide for a more comprehensive model for estimating emotional response. Additionally, the insights from this research have potential applications in game development, such as adaptive music systems for video games, procedural content generation, and therapeutic virtual reality experiences.

Information

Lärosäte / institution
Stockholms universitet/Institutionen för data- och systemvetenskap
Publiceringsdatum
2025
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.