Uppsats

Evaluating the Cognitive Plausibility of Transformer-Based Models: Predicting Articulation Rate in Read Speech from Surprisal Estimates

Master-uppsats

Uppsala universitet/Institutionen för lingvistik och filologi

Publicerad: 2025

Språk: Engelska

Sammanfattning

Models of human communication posit that the transmission of information between speakers and listeners is constrained by a noisy channel, from which listeners must identify the intended message despite uncertainty in the signal. Overcoming uncertainty depends on the predictability and informativeness of the linguistic units in the message, which influence the cognitive effort required for language processing and comprehension. Modeling cognitive mechanisms of human language processing involves quantifying this cognitive effort, commonly measured by surprisal. Surprisal is a measure of the amount of information carried by a linguistic unit (e.g., a word), which is inversely related to its predictability: highly predictable words have low information content and unexpected words carry more information. Research has shown a negative relationship between surprisal and speech tempo—such as articulation rate (AR)—as reflected in human behavioral data, including reading times. Transformer-based Large Language Models (LLMs) like GPT-2 produce surprisal estimates that are strongly predictive of reading times in humans, effectively capturing this negative relationship. However, recent analyses have observed a positive relationship between the Psychometric Predictive Power (PPP) of larger GPT-2 models and perplexity (PPL)—compared to smaller GPT-2 variants—suggesting that larger models poorly approximate human reading behavior. However, it remains unexplored whether surprisal estimates from BERT can better predict AR in read speech than GPT-2. The aim of this work is to evaluate which of GPT-2 and BERT are better cognitive models of language processing in humans. We employ Generalized Linear Models (GLMs) to model AR as a function of surprisal estimates from both models, while controlling for word position, word length, and final lengthening. The results suggest that BERT is a more cognitively plausible model of human language processing than GPT-2, based on two observations. Firstly, BERT models trained with smaller context window sizes obtained surprisal estimates that led to improved PPP. Secondly, the smallest BERT model—based on a context window size of 128 tokens—achieved the lowest Akaike's Information Criterion (AIC) score, suggesting a better fit to the observed data. We argue that our results have broader implications for cognitive modeling, especially the possibility of BERT to capture different psycholinguistic processes in humans than GPT-2—language planning, retrospective simulation, and the interaction between prediction and comprehension. Our results question prior assumptions that BERT is a cognitively implausible model.

Information

Författare
Giachi, Matteo
Lärosäte / institution
Uppsala universitet/Institutionen för lingvistik och filologi
Publiceringsdatum
2025
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.