Uppsats

Developing a Framework for Semantic Disagreement Interruption Turn-taking Model

Master-uppsats

Blekinge Tekniska Högskola/Institutionen för datavetenskap

Publicerad: 2025

Språk: Engelska

Sammanfattning

Background: Spoken Dialogue Systems increasingly strive to exhibit natural and context-aware conversational behaviour. While existing turn-taking models have achieved progress in predicting turn-shifts and backchannels using prosodic or lexical cues, the ability to strategically interrupt based on semantic understanding, such as when a user expresses disagreement or factual inconsistency, remains largely unexplored. Objectives: This thesis investigates the development of a real-time interruption framework capable of detecting semantic disagreement between a user’s spoken input and a trusted reference statement. The objective is to enable a conversational agent to make contextually appropriate interruption turn-taking decisions by modelling both the stance relationship and the pragmatic conditions under which interruptions are justified. Methods: A custom dataset was constructed by adapting the Fake News Challenge (FNC-1) corpus into a conversational, speech-like format encompassing eight semantically nuanced classes. The framework recommends suitable components for real-time operation, including an automatic speech recognition (ASR) architecture for transcription, scalable reference retrieval through similarity search, and a pre-trained encoder-only transformer for stance classification. Model performance was evaluated using accuracy, macro F1, and an F1-score on the class of interest. Results: The proposed dataset comprises approximately 1,400 labelled text pairs across eight semantically nuanced classes. The Transformer Transducer architecture was identified as an appropriate ASR model for real-time transcription. The DistilBERT encoder-only transformer was fine-tuned and evaluated under varying training configurations, achieving a best accuracy of 0.66, macro F1 of 0.51, and 0.48 F1 score on the class of interest. Precision–recall analysis further revealed that model responsiveness could be fine-tuned via threshold selection to balance conservative and assertive interruption strategies. The framework also demonstrated scalability through similarity search and responsiveness through parallel ASR-classification processing. Conclusions: The study highlights both the feasibility and challenges of modelling interruption behaviour driven by semantic disagreement. Although current performance remains moderate due to dataset size and class imbalance, the proposed framework establishes a foundation for real-time, semantically grounded interruption modelling. Future directions include integrating spoken language models, reasoning-based decision mechanisms, and empirical identification of temporal cues for interruption timing to develop more socially aware and human-like conversational agents.

Information

Lärosäte / institution
Blekinge Tekniska Högskola/Institutionen för datavetenskap
Publiceringsdatum
2025
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.