Uppsats
Privacy-Preserving Semantic Candidate Matching : Using Vector Retrieval and LLM-Based Analysis
Yrkesexamen på grundnivå
Karlstads universitet/Institutionen för matematik och datavetenskap (from 2013)
Publicerad: 2026
Språk: Engelska
Sammanfattning
Modern recruitment systems increasingly rely on processing unstructured textual data such as Curriculum Vitae (CVs) and Job Descriptions (JDs). At the same time, the use of Artificial Intelligence (AI) on internal enterprise documents introduces challenges related to privacy and regulatory compliance due to the presence of Personally Identifiable Information (PII). This thesis investigates how anonymization, semantic vector retrieval, and Retrieval-Augmented Generation (RAG) can be combined within a privacy-preserving candidate matching system. A functional prototype was designed and implemented using document extraction, Named Entity Recognition (NER)-based anonymization, embedding generation, vector-based retrieval, and Large Language Model (LLM)-based analysis. The results demonstrate that semantic retrieval could identify relevant candidates even when competencies were expressed using different terminology across documents. The anonymization pipeline successfully removed sensitive information while preserving sufficient semantic content for retrieval and analysis. Furthermore, the system generated structured candidate evaluations based on retrieved contextual information. The implemented prototype demonstrates the technical feasibility of combining anonymization, vector retrieval, and LLM-based analysis within a unified workflow.
Information
- Författare
- Gustafsson, Emil, Silakhory, Pedram
- Lärosäte / institution
- Karlstads universitet/Institutionen för matematik och datavetenskap (from 2013)
- Publiceringsdatum
- 2026
- Uppsatstyp
- Yrkesexamen på grundnivå
- Språk
- Engelska