Uppsats
Developing a Framework for Semantic Disagreement Interruption Turn-taking Model
Master-uppsats
Blekinge Tekniska Högskola/Institutionen för datavetenskap
Publicerad: 2025
Språk: Engelska
Sammanfattning
Background: Spoken Dialogue Systems increasingly strive to exhibit natural and context-aware conversational behaviour. While existing turn-taking models have achieved progress in predicting turn-shifts and backchannels using prosodic or lexical cues, the ability to strategically interrupt based on semantic understanding, such as when a user expresses disagreement or factual inconsistency, remains largely unexplored. Objectives: This thesis investigates the development of a real-time interruption framework capable of detecting semantic disagreement between a user’s spoken input and a trusted reference statement. The objective is to enable a conversational agent to make contextually appropriate interruption turn-taking decisions by modelling both the stance relationship and the pragmatic conditions under which interruptions are justified. Methods: A custom dataset was constructed by adapting the Fake News Challenge (FNC-1) corpus into a conversational, speech-like format encompassing eight semantically nuanced classes. The framework recommends suitable components for real-time operation, including an automatic speech recognition (ASR) architecture for transcription, scalable reference retrieval through similarity search, and a pre-trained encoder-only transformer for stance classification. Model performance was evaluated using accuracy, macro F1, and an F1-score on the class of interest. Results: The proposed dataset comprises approximately 1,400 labelled text pairs across eight semantically nuanced classes. The Transformer Transducer architecture was identified as an appropriate ASR model for real-time transcription. The DistilBERT encoder-only transformer was fine-tuned and evaluated under varying training configurations, achieving a best accuracy of 0.66, macro F1 of 0.51, and 0.48 F1 score on the class of interest. Precision–recall analysis further revealed that model responsiveness could be fine-tuned via threshold selection to balance conservative and assertive interruption strategies. The framework also demonstrated scalability through similarity search and responsiveness through parallel ASR-classification processing. Conclusions: The study highlights both the feasibility and challenges of modelling interruption behaviour driven by semantic disagreement. Although current performance remains moderate due to dataset size and class imbalance, the proposed framework establishes a foundation for real-time, semantically grounded interruption modelling. Future directions include integrating spoken language models, reasoning-based decision mechanisms, and empirical identification of temporal cues for interruption timing to develop more socially aware and human-like conversational agents.
Information
- Författare
- Kandula, Shiva Chandra
- Lärosäte / institution
- Blekinge Tekniska Högskola/Institutionen för datavetenskap
- Publiceringsdatum
- 2025
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Umeå universitet/Institutionen för informatik
Alaniemi, Peter, Bergman, Elias
Publicerad: 2025
Master-uppsats, Linköpings universitet/Människocentrerade system
He, Zhuoyou
Publicerad: 2025
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Tinke, Pjotr
Publicerad: 2025
Master-uppsats, Uppsala universitet/Institutionen för informatik och media
Lundmark, Rebecka, Berggren, Daniel
Publicerad: 2025
Master-uppsats, Uppsala universitet/Institutionen för informatik och media
Mahanta, Uddipta
Publicerad: 2025
Master-uppsats, Högskolan i Halmstad/Akademin för informationsteknologi
Ottakath, Jayitha
Publicerad: 2025