Uppsats

Identifying and Explaining Anomalous Weeks in Issues from Customer Services : Large language models and classical approaches in synergy

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2025

Språk: Engelska

Sammanfattning

This thesis investigates methods for identifying and explaining anomalous weeks in customer service data from Stockholm’s public transport agency. The work is divided into three parts: detecting anomalies in time series, automatically subdividing issue categories into more specific subcategories and generating explanations for anomalous periods. The approach combines classical data science techniques, such as clustering and statistical anomaly detection, with state-of-the-art large language models (LLMs) applied to a dataset in Swedish. Anomaly detection was conducted using four models: two established tools (Prophet and SARIMA) and two custom methods based on mean and standard deviation thresholds within sliding windows. The evaluation showed that no model consistently outperformed the others across all time series, all exhibiting both false positives and false negatives. However, the most promising results were obtained from Prophet and a normalised sliding window approach that accounts for passenger volume. To split a broad customer service category into more informative subcategories, a clustering-based method using text-embedding-3-large was compared to zero-shot classification using GPT-4o and GPT-4o-mini. The LLM-based approach, particularly GPT-4o, produced more coherent and distinct category suggestions and achieved higher classification accuracy in a labelled test dataset. Explanations for anomalous weeks were generated by prompting LLMs with customer service issues from the anomalous period and the preceding month. Four models were tested: GPT-4o, GPT-4o-mini, o1 and o3-mini. Explanations were evaluated using a four-part rubric: accuracy, insight, structure & clarity and brevity. All models produced useful summaries, but exaggeration of minor patterns was a recurring issue. The reasoning models, o1 and o3-mini, were generally rated higher than the GPT models. Overall, the results highlight the potential of combining generative AI with classical methods to support scalable, data-driven analysis of customer feedback in public transport operations.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.