Uppsats
Design and Evaluation of an AI-Driven Task-to-Code Feedback Loop in CI/CD Pipelines
Kandidat-uppsats
Blekinge Tekniska Högskola/Institutionen för programvaruteknik
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Background. CI/CD pipelines shorten feedback cycles and enable early defect detection, but the path from a task description to a validated pull request still relies on manual coordination despite growing AI adoption in software development. Objectives. We design and evaluate ACID Bot (AI in CI/CD), a bounded three-agent pipeline (Enhancer, Solver, Publisher) built on the Claude Code CLI that retrieves bug-fix tasks from ClickUp via the Model Context Protocol, generates code and tests, and opens a pull request for human review. We assess end-to-end feasibility, comparative performance, and operational constraints in an industrial deployment. Methods. We conduct a bounded industrial case study at Iquest AB across seven real Java bug-fix tasks (single-service backend defects, excluding UI, database migration, security, and multi-service tasks), comparing available manual, AI-assisted, and autonomous conditions. Historical baseline timings are taken from ClickUp time tracking, contextual observations and run metadata from a structured questionnaire completed by the Iquest AB contact, and merge outcomes from GitHub pull-request artifacts. Results. ACID Bot reached pull request creation for all seven tasks without manual code changes (median 7 minutes, versus 60 manual and 450 AI-assisted), but only four of seven pull requests were merged; T5, T6, and T7 were not, because fixes were partial, misdirected, or symptom-level. The Solver's self-assessed Oracle scores ranged 90-97 (median 95) but did not predict merge outcome and are reported as an agent self-confidence signal rather than a quality measure. Three failure patterns emerged: mixed-intent task descriptions, incomplete root-cause resolution, and cross-cutting partial coverage. Conclusions. Within this bounded industrial case, ACID Bot created pull requests more quickly on the directly comparable non-T7 AI-assisted task pairs, while only four of seven generated pull requests were accepted for merge. Task-description clarity and available diagnostic context emerged as important observed factors, and mandatory human review identified the three insufficient fixes - each of which had passed automated tests and received a high self-assessed Oracle score.
Information
- Författare
- Ilhomov, Abduvohid, Matar, Karam
- Lärosäte / institution
- Blekinge Tekniska Högskola/Institutionen för programvaruteknik
- Publiceringsdatum
- 2026
- Uppsatstyp
- Kandidat-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Yrkesexamen på grundnivå, Högskolan i Halmstad/Akademin för informationsteknologi
Hamedy, Faraj, Mouselli, Alaa
Publicerad: 2025