Uppsats

Consolidation of Software Documentation Using LLMs : Investigating the Usefulness of Agentic RAG Systems When Creating Structured Documentation

Yrkesexamen på avancerad nivå

Blekinge Tekniska Högskola/Institutionen för programvaruteknik

Publicerad: 2026

Språk: Engelska

Sammanfattning

Background. Software documentation in commercial renovation projects is scattered across heterogeneous sources — wikis, issue trackers, slide decks, spreadsheets — and accumulates as documentation debt over a project's lifetime. Producing structured artefacts from this corpus is currently a manual task for project managers. Large Language Models (LLM) and Retrieval-Augmented Generation (RAG) suggest a possible automation path, but their effectiveness in industrial settings remains unclear. Objectives. This thesis investigates (i) what documentation consolidation looks like in practice in commercial software renovation projects and what requirements practitioners place on a supporting tool, and (ii) how, and how effectively, generative AI can support this consolidation when producing structured arc42 architecture documents. Methods. Following Design Science Research, we conducted seven semi-structured interviews with project managers at our industry partner Itestra to elicit requirements (RQ1). We then designed a prototype pipeline combining hybrid BM25 and dense retrieval with reciprocal-rank fusion, contextual chunking, link-based cross-chunk reference following, hierarchical topic mapping, and a sequential plan–draft–critique–revise agent (RQ2). The prototype was evaluated through six practitioner interviews, RAGAS faithfulness measurement, and a LLM-as-judge comparison against a baseline agent (Claude Code) with the same underlying LLM and direct file access, but without our retrieval and preprocessing pipeline. Literature engagement is non-systematic and partly LLM-assisted; coverage claims are scoped accordingly. Results. RQ1 surfaced four cross-cutting concerns: poor findability, code-as-truth epistemics, audience fragmentation, and compliance-versus-use tension. RQ2 produced a coherent arc42 artifact but did not reliably outperform the baseline (16 of 28 (judge × section) cells, covering 3 of 7 arc42 sections, with Correctness excluded from the LLM-as-judge comparison), and faithfulness scores were low across most sections. Practitioners rejected the document-output paradigm itself, instead proposing diagnostic or query-driven tools. Conclusions. Generative AI can support documentation consolidation in a functional sense, but a single static artifact attempting to serve all audiences systematically underserves each. Three structural constraints follow: consolidation relocates rather than dissolves the volume problem; source-hierarchy alignment bounds output trustworthiness more strongly than retrieval architecture does; and audience scoping must be treated as a primary design parameter rather than an emergent property of the consolidation process. These findings reframe the design space for AI documentation tools toward audience-scoped outputs and diagnostic operation on the source corpus.

Information

Lärosäte / institution
Blekinge Tekniska Högskola/Institutionen för programvaruteknik
Publiceringsdatum
2026
Uppsatstyp
Yrkesexamen på avancerad nivå
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.