Uppsats

Evaluating Generative Language Models in Japanese Using Question-Answering Templates : A Synthetic Data Generation Approach

Master-uppsats

Uppsala universitet/Institutionen för lingvistik och filologi

Publicerad: 2024

Språk: Engelska

Sammanfattning

In the last five years, transformer-based large language models have not only made tremendous advancements in fluent and accurate text generation, but also show unparalleled performance as zero- or few-shot shot learners of numerous downstream tasks. As their arsenal continues to expand, so must our evaluation methods. This thesis project contributes to an existing linguistically motivated template-based question-answering behavior testing framework, the Multilingual Mor-phological Checklist, by developing Japanese templates for a number of logical reasoning tasks. Furthermore, we add politeness as a new linguistic dimension within the framework. The Japanese question-answering prompts generated by our templates are evaluated on three pre-trained transformer-based multilingual generative language models alongside English, German and Finnish prompts. The model performance is evaluated in three experimental settings: one zero-shot experiment and two one-shot experiments. Using a number of automatic partial string-matching strategies, we find that all models achieve reasonable accuracy on the Japanese prompts for a majority of the reasoning tasks. The accuracy on Japanese is, unsurprisingly, worse than for English on average. Moreover, the best performance is observed in the zero-shot setting, suggesting that the introduction of additional elements in prompts may confuse the model more than the example helps clarify the task format. Additionally, we evaluate the reliability of the automated string-matching strategies, comparing them to a sample of human annotation. This reveals that the English scores align closer to the human scores than Japanese scores do, suggesting that different languages may benefit from different automated scoring approaches. Lastly, we find that although the performance across the politeness dimensionis quite scattered, formal Japanese prompts sometimes outperform the polite and familiar language variants. On average, however, the formal language prompts result in poorer performance, which aligns with our initial expectations.

Information

Författare
Norrman, Victor
Lärosäte / institution
Uppsala universitet/Institutionen för lingvistik och filologi
Publiceringsdatum
2024
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.