Uppsats
Reducing Carbon Emissions in k-Means Clustering Using Representative Subsets
Master-uppsats
Linköpings universitet/Institutionen för datavetenskap
Publicerad: 2025
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
The growing scale of data and the complexities of machine learning models have led to significant energy-related carbon emissions, which threaten ecosystems and communities worldwide, prompting the search for greener alternatives. Research regarding the energy consumption in machine learning has been biased towards large NLP models such as BERT and GPT, leaving a gap in understanding the sustainability implications of other fundamental algorithms. This thesis investigates how representative subsets can reduce carbon emissions on the example of the k-means algorithm without significantly compromising performance. Three well-known subset constructions are evaluated: Uniform sampling, the Lightweight Coreset and the Strong Coreset for Bregman clustering, along with two novel approaches designed to improve the construction time of the Strong Coreset: one replaces the k-means++ initialization in the Strong Coreset with the AFK-MC2 approximation, while the other reduces the dimension of the dataset with the Clarkson-Woodruff transform before the coreset construction. The experiments were run on 12 datasets of different sizes and complexities, testing two k values per dataset and six subset sizes. Performance was measured using Sum of Squared Distance (SSD) and Davis-Bouldin Score (DBS) along with measurements of execution time and energy consumption. Each combination of dataset, k value and subset size was also tested with the Kruskal-Wallis test to determine whether a significant difference existed between the methods for any of the measured metrics. When a Kruskal-Wallis test indicated a significant result, the Dunn’s-test was applied to identify which specific methods differed from each other. Results demonstrated that while in most cases there is a significantly different distribution between running k-means on the full dataset and using a coreset regarding the performance metrics SSD and DBS, the practical magnitude of these differences is minimal. A subset size of 5 percent typically remains less than 5 percent away of the same performance of running on the full dataset across the tested datasets and k values used in this thesis for the AFK-MC2, Lightweight, Strong Coresets. All combinations of datasets, subset sizes and k values have a significantly different distribution in execution time and energy consumption compared to the full data, with all of the results from representative subsets having a much lower median. A strong correlation between time and energy reduction was observed for every dataset with the lowest value of 0.966. However, the slope of this relationship varied depending on the dataset, k and sometimes also between representative subset methods. The Lightweight Coreset often outperformed the Strong Coresets in efficiency except in very specific configurations unlikely to occur in practice. Dimension reduction via the Clarkson-Woodruff Transform showed promising time and energy improvements for the Strong Coreset with benefits increasing for larger k values and subset sizes. Additional experiments on parallelization for the k-means algorithm with a k of 50 on small subsets of 397 and 1983 observations and 58 features, showed a significant reduction in energy consumption when running k-means in parallel. These findings suggest that representative subsets combined with algorithmic optimization offer an effective approach to reduce the carbon emissions of k-means clustering while maintaining high performance levels.
Information
- Författare
- Hedström, Johannes
- Lärosäte / institution
- Linköpings universitet/Institutionen för datavetenskap
- Publiceringsdatum
- 2025
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Försvarshögskolan
Hellqvist, Theodor
Publicerad: 2026
Kandidat-uppsats, Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Ksouri, Zied, Sai, Adam
Publicerad: 2026
Master-uppsats, Göteborgs universitet/Graduate School
Mikhail, Maryana
Publicerad: 2026-07-09
Master-uppsats, Göteborgs universitet/Graduate School
Melin, Gustav
Publicerad: 2026-07-09
Master-uppsats, Göteborgs universitet/Graduate School
Enges, Emil, Lundgren, Olle
Publicerad: 2026-07-02
Master-uppsats, Luleå tekniska universitet/Institutionen för ekonomi, teknik, konst och samhälle
Krawe, Manoj Nalaka Sanjeewa, Peiris, Pattiyage Jayani Yasoja
Publicerad: 2026