Uppsats

AI Generated code documentation : A study on Large Language Model utilization for code documentation using limited data

Yrkesexamen på avancerad nivå

Umeå universitet/Institutionen för tillämpad fysik och elektronik

Publicerad: 2025

Språk: Engelska

Sammanfattning

Large language models (LLMs) are a class of computational models that have the ability to understand and generate human language. This study aims to investigate and evaluate the performance of large language models in generating code documentation using limited training data, by comparing different techniques such as Retrieval Augmented Generation (RAG) and fine-tuning, to determine the effectiveness in enhancing the LLM performance with specific data. The evaluation uses a limited dataset of PHP code, using a number of different large language models. The models that are used for the testing are a quantized version of the Gemma 2 model with 9 billion parameters and a quantized version of the Codellama model with 7 billion parameters. Snippets of code unseen documents were provided as input for the models to generate corresponding documentation. The documentation were then evaluated and scored by developers. The results show that utilizing prompt engineering combined with fine-tuning gives the best result out of the tested models and strategies. The fine-tuned Gemma 2 models final score was 8.59 out of 10, which was the highest, while the Gemma 2 with RAG scored 6.41 out of 10, this being the lowest.

Information

Författare
Salih, Miran
Lärosäte / institution
Umeå universitet/Institutionen för tillämpad fysik och elektronik
Publiceringsdatum
2025
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.