Uppsats

Exploring an AI Agentic Workflow for Solving Challenging Coding Problems : An Evaluation of a Large Language Model Based Multi-Agent System

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2025

Språk: Engelska

Sammanfattning

Generative Artificial Intelligence (Gen AI) models have become very popular, especially after the public release of ChatGPT in late 2022. The demonstrated capabilities of these models have led researchers to use them to build agentic workflows, including Multi-Agent Systems (MAS), and explore their potential. Previous studies that investigate these systems’ coding abilities have shown promising results. However, these studies have mainly used coding benchmarks consisting of simple coding problems, leaving a gap in understanding coding ability of agentic workflows when facing coding problems of greater difficulty. This study tries to fill this gap by developing a MAS and evaluating its coding ability against 120 selected coding problems taken from Kattis, with the difficulty levels ranging from easy to hard. In the experiment, both the developed MAS (configured with Llama 3-70b) and a single Large Language Model (LLM) (Llama 3-70b) were tested on the selected coding problems using Kattis’ assessment system. The experiment showed that the MAS was significantly better than the single LLM when zero- shot prompting was used, increasing acceptance rates by 6.7% and decreasing failure rates by 11.7%. Overall, the results of this study suggest that using a designed LLM-based MAS offers a significant performance boost compared to a single LLM.

Information

Författare
Berlin, Claudia
Lärosäte / institution
KTH/Skolan för elektroteknik och datavetenskap (EECS)
Publiceringsdatum
2025
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.