Uppsats

Design and Evaluation of AI-Assisted Information Extraction Systems for Enterprise Environments

Master-uppsats

Uppsala universitet/Institutionen för informatik och media

Publicerad: 2026

Språk: Engelska

Sammanfattning

Enterprise organisations rely heavily on large collections of unstructured and semi-structured documents to support engineering, compliance, and operational activities. Extracting structured information from these documents is often performed manually, making the process time-consuming, difficult to scale, and dependent on individual interpretation. Recent advances in artificial intelligence (AI) and large language models (LLMs) have created new opportunities for automating information extraction from complex enterprise documents. However, practical deployment remains challenging due to heterogeneous document structures, contextual ambiguity, and the need for traceable and consistent outputs. This thesis investigates how different AI-assisted extraction pipeline architectures behave when applied to enterprise product requirement documents. The study focuses on the extraction of references to external standards and associated metadata from heterogeneous Word and Excel documents obtained through collaboration with Ericsson. Using a Design Science Research (DSR) methodology, three alternative extraction pipelines were designed and evaluated: a custom Multi-Model Pipeline (MMP), an Agent Full-Document Pipeline (AFP), and an Agent Chunked Pipeline (ACP). The pipelines represent different design choices related to preprocessing, chunking, orchestration, and model coordination. The evaluation compares the pipelines in terms of extraction coverage, metadata completeness, consistency, traceability, cross-pipeline agreement, and processing behaviour. Since comprehensive validated ground-truth datasets were not available, the study focuses primarily on comparative behavioural analysis rather than definitive accuracy benchmarking. The results show that architectural and preprocessing decisions substantially influence extraction behaviour. The multi-model pipeline achieved the broadest extraction coverage, while structured preprocessing and chunking improved output consistency, metadata completeness, and traceability. The findings also indicate that document size and structural complexity increase variability across both standard extraction and contextual metadata identification. The thesis contributes empirical insights into the trade-offs associated with AI-assisted information extraction in enterprise environments. In particular, the study highlights the importance of preprocessing, orchestration strategy, and workflow integration when designing AI-assisted extraction systems intended to support organisational decision-making processes.

Information

Författare
Nair, Aditya
Lärosäte / institution
Uppsala universitet/Institutionen för informatik och media
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.