Uppsats
Evaluating Generative Language Models in Japanese Using Question-Answering Templates : A Synthetic Data Generation Approach
Master-uppsats
Uppsala universitet/Institutionen för lingvistik och filologi
Publicerad: 2024
Språk: Engelska
Sammanfattning
In the last five years, transformer-based large language models have not only made tremendous advancements in fluent and accurate text generation, but also show unparalleled performance as zero- or few-shot shot learners of numerous downstream tasks. As their arsenal continues to expand, so must our evaluation methods. This thesis project contributes to an existing linguistically motivated template-based question-answering behavior testing framework, the Multilingual Mor-phological Checklist, by developing Japanese templates for a number of logical reasoning tasks. Furthermore, we add politeness as a new linguistic dimension within the framework. The Japanese question-answering prompts generated by our templates are evaluated on three pre-trained transformer-based multilingual generative language models alongside English, German and Finnish prompts. The model performance is evaluated in three experimental settings: one zero-shot experiment and two one-shot experiments. Using a number of automatic partial string-matching strategies, we find that all models achieve reasonable accuracy on the Japanese prompts for a majority of the reasoning tasks. The accuracy on Japanese is, unsurprisingly, worse than for English on average. Moreover, the best performance is observed in the zero-shot setting, suggesting that the introduction of additional elements in prompts may confuse the model more than the example helps clarify the task format. Additionally, we evaluate the reliability of the automated string-matching strategies, comparing them to a sample of human annotation. This reveals that the English scores align closer to the human scores than Japanese scores do, suggesting that different languages may benefit from different automated scoring approaches. Lastly, we find that although the performance across the politeness dimensionis quite scattered, formal Japanese prompts sometimes outperform the polite and familiar language variants. On average, however, the formal language prompts result in poorer performance, which aligns with our initial expectations.
Information
- Författare
- Norrman, Victor
- Lärosäte / institution
- Uppsala universitet/Institutionen för lingvistik och filologi
- Publiceringsdatum
- 2024
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Karlstads universitet/Institutionen för matematik och datavetenskap (from 2013)
Florean, Alexander
Publicerad: 2024
Master-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Shiva Olin, Harald
Publicerad: 2025
Master-uppsats, KTH/Fysik
Öqvist, Linda
Publicerad: 2026
Kandidat-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Abou Hachem, Rola
Publicerad: 2024
Kandidat-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Qutbuddin Habib, Hiba, Belhajali, Hibatallah
Publicerad: 2024
Kandidat-uppsats, KTH/Skolan för elektroteknik och datavetenskap (EECS)
Amani Rudäng, Elias, Alhallak, Natali
Publicerad: 2025