Uppsats

Novel LLM Benchmark for Unit Test Generation

Master-uppsats

Linköpings universitet/Institutionen för datavetenskap

Publicerad: 2026

Språk: Engelska

Sammanfattning

Software testing has always been an important part of software development. As long as software testing has existed, there have also been attempts to automate it. Nowadays, with the rise of Large Language Models (LLMs) in our daily lives, many attempts have been made to use them for test generation, especially for unit tests. Despite the growing demand, there is still no benchmark specialized for industrial use that evaluates unit test generation in Java. This work aims to address this gap by creating an automated benchmark for unit test generation, which uses dynamic prompting techniques and five metrics to evaluate the performance of each LLM. It covers three difficulty levels (Simple, Complex, and Industry), which makes it applicable in an industry setting. This benchmark was tested on four LLMs (DeepSeek-R1, Mistral-Nemo, Llama3, and Qwen2.5). The results show the effectiveness of dynamic prompting compared to static prompt templates. We also observe that, while models show promising results in the Simple and Complex categories, the results in the Industry category still have significant room for improvement. Furthermore, the results demonstrate that the size of a model is not always correlated to its performance in this task, which highlights the importance of benchmarking.

Information

Författare
Hashemi, Afshan
Lärosäte / institution
Linköpings universitet/Institutionen för datavetenskap
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.