Uppsats

Developing a Method for Investigating the Population of Vocalists Heard in AI-Generated Music

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2025

Språk: Engelska

Sammanfattning

Despite the increase of music generated by Artificial Intelligence (AI) tools such as Suno and Udio, little research has been done on the songs they create. Previous research has largely focused on techniques for AI music detection, while the potential biases and patterns in the vocals the models generate have been left unanalyzed. This thesis aims to develop a pipeline that allows for population analysis of singing voices present in music generated by Suno and Udio. In order to accomplish this, we investigate two types of methods. One approach uses Mel-Frequency Cepstrum Coefficients (MFCCs) for feature extraction together with Gaussian Mixture Models (GMMs) to model vocal characteristics. The other approach uses deep learning models to extract features directly, with two speaker recognition models and one singing voice representation model. We evaluate both methods through testing with real songs that we process using source separation and silence removal. Based on the initial test results we then apply the most reliable model — the singer representation model — to the dataset of AI singers and use K-medoids clustering as well as Uniform Manifold Approximation and Projection (UMAP) dimensionality reduction to examine the data. The results from the singer representation model showed only limited quantitative success with the K-medoids clustering, while qualitative testing on real songs suggests the approach is somewhat successful. The UMAP projection showed quite distinct separation between the Suno and Udio songs, as well as between female and male vocals. Issues with distortion or incorrectly converted audio files in the datasets we used were discovered, which likely had significant impact on the clustering and visualization results. An implementation error in our UMAP usage of the MFCC approach initially showed it as being much less reliable than further testing seemed to indicate, making this method possibly interesting for more thorough testing. We motivate further research into how the voice characteristics of the generated vocals relate to their textual prompts and vocals of real artists.

Information

Författare
Nylén, Tore
Lärosäte / institution
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publiceringsdatum
2025
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.