Uppsats

From Code to Competence : Assessing LLMs’ Ability to Generate and Improve Code

Kandidat-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2025

Språk: Engelska

Sammanfattning

Artificial Intelligence (AI), and in particular large language models (LLMs), has over the past years emerged as one of the most transformative technologies, igniting a paradigm shift across industries by redefining how tasks might be approached, handled, and completed. One prominent application of AI is code generation, where LLMs are increasingly employed by developers in industry and students in higher education in performing programming tasks. As these models evolve, understanding their true ‘competence’ in generating functional and efficient code has become a topic of growing interest, particularly within educational settings, where concerns about academic integrity and skill development have emerged. This study builds upon existing research exploring LLMs’ capabilities, particularly in the context of code generation and prompt engineering, and seeks to fill gaps of knowledge currently present due to rapid advancements in the field. In phase 1 of the study, base prompting was used with two of the most widely used LLMs currently – ChatGPT-4o, Deepseek-V3 – across 150 Javascript programming problems from LeetCode, evenly categorized as easy, medium, or hard. In phase 2, the failed tasks were reattempted with three prompting strategies – feedback prompting, multi-step conversational prompting, and 1-shot example prompting – that have, in past research, proved to have a significant impact on LLMs’ performance. In the results observed, both models reached a 100% completion rate for easy problems, suggesting that these two LLMs have reached full competence within this category. For problems categorized as medium difficulty, both models completed over 80%, and for problems categorized as hard, over 50%. In the second phase, feedback prompting proved to be the prompting strategy with the greatest impact, regardless of which LLM and difficulty level it was being used with While there was some linearity observed between different parameters for the various prompting strategies, certain anti correlation was observed. Thereby, a conclusive verdict that different prompting strategies have a varying impact depending on what LLM it is used with, could not be ruled out. A large positive correlation was, however, found in relation to programming task difficulty, where the prompting strategies were observed to have a greater impact when used for tasks of a higher level of difficulty.

Information

Lärosäte / institution
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publiceringsdatum
2025
Uppsatstyp
Kandidat-uppsats
Språk
Engelska