Uppsats

Now or Later. Classifying Future- and Present-Oriented Speech Acts in Organizational Communication : Levaraging Large Language Models in automatic dataset production

Kandidat-uppsats

Linköpings universitet/Institutionen för datavetenskap

Publicerad: 2024

Språk: Engelska

Sammanfattning

As companies transition to digital platforms the imperatives of preventative innovation that often do not yield immediate economic benefit (Jönsson et al., 2024), such as ensuring data integrity and security often contend with innovations aimed at immediate economic benefits (Jönsson et al., 2024). In this environment effective organisational communication becomes a means to alleviate the imbalance between the immediate economic benefits of innovations aimed at addressing current demands and the delayed economic gratification of preventative innovation aimed at addressing future needs (Jönsson et al., 2024).Ancillary to the research conducted by Jönsson et al. (2024) regarding the economic recognition of organisations adoption of information security standards ISO/IEC 27001,this thesis seeks to provide methods and tools to aid deeper analysis of the linguistic mechanisms through which, adherence of such standards are articulated and disseminated. Specifically, through the classification of the Speech Acts employed in organisational discourse regarding ISO/IEC 27001 standards as either expressed intentions (future-oriented) or statements (present-oriented).Given the natural language processing task of classifying Speech Acts, previous work often relies on large amounts of manually annotated data occasionally combined with hand crafted features in order to train classification models. Taking inspiration from Sun et al’s (2023) work in utilising LLMs for text classification using Clue and Reasoning Prompting (CARP) and Retrieval Augmented Prompting (RAG), this thesis proposes to circumvent the need for large amounts of manually annotated data by leveraging the capabilities of LLMs, in the task of Speech Act categorisation. Initially, a model was selected based on its performance with a small, manually annotated dataset. This was followed by hyperparameter tuning to optimise the model. A dataset was then created for demonstration retrieval, starting with the model categorising sections of the corpus, which were subsequently manually annotated. Clues and reasoning were added to each entry, following the approach described in Sun et al. (2023). The process continued with fine-tuning a XLM-RoBERTa model on this data for classification-sensitive demonstration retrieval, aligning with Sun et al’s methods (2023). CARP prompting (Sun et al., 2023) was used to annotate approximately 56% of the corpus. Finally, another XLM-RoBERTa model was trained on this automatically annotated data and evaluated using a manually annotated test dataset. The classifier model achieved an accuracy score of 73% with a weighted F1-score of 0.75 on the test dataset. The automatically annotated dataset comprised 138 052 entries excluding instances for which no annotation was found. The results of the thesis show that utilising LLMs for automatic dataset generation can be a viable alternative or supplement to manual annotation.

Information

Författare
Grattan, Marcus
Lärosäte / institution
Linköpings universitet/Institutionen för datavetenskap
Publiceringsdatum
2024
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.