Uppsats

Evaluating the impact of chunking strategies and embedding models on retrieval performance in naive RAG

Kandidat-uppsats

Högskolan i Skövde/Institutionen för informationsteknologi

Publicerad: 2026

Språk: Engelska

Sammanfattning

Large Language Models (LLMs) have advanced Natural Language Processing (NLP) but remain prone to hallucinations, particularly in domain-specific applications such as legal text processing. Retrieval-Augmented Generation (RAG) addresses this limitation by grounding responses in retrieved documents. This study investigates the main and interaction effects of chunking strategies and embedding models on retrieval performance in a Naive RAG pipeline using a GDPR-based legal corpus. A controlled factorial experiment evaluated recursive, token-based and sentence-based chunking in combination with the OpenAI text-embedding- 3-small and text-embedding-3-large models using Recall@k, Mean Reciprocal Rank (MRR) and normalized Discounted Cumulative Gain (nDCG@k). The results show that both the chunking strategy and the embedding model have statistically significant effects on retrieval performance, with the chunking strategy being the dominant factor. A statistically significant interaction effect was also detected, although the effect sizes for both the main and interaction effects were small. Sentence-based chunking achieved the strongest overall performance, while token-based chunking produced the highest MRR. Inferential analyses indicate that these differences should be interpreted with caution, given their limited practical significance.

Information

Lärosäte / institution
Högskolan i Skövde/Institutionen för informationsteknologi
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.