Uppsats
Automatic Anonymization of Personal Identifiable Information in Images of Unstructured Documents - In relation to Swedish-specific PII
Kandidat-uppsats
Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Publicerad: 2026
Språk: Engelska
Sammanfattning
In today's digital landscape, the widespread sharing of unstructured documents such as receipts and invoices presents a significant privacy risk, particularly concerning region-specific sensitive data. This bachelor's degree project investigates the automated anonymization of Swedish-specific Personal Identifiable Information (PII) within unstructured document images, addressing the knowledge gap left by existing systems trained predominantly on English-language and US-centric datasets. To tackle this, a prototype pipeline was developed using a Design Science methodology, integrating various Optical Character Recognition (OCR) tools (EasyOCR, Mindee DocTR, PaddleOCR, and Tesseract) and BERT-based/-like Named Entity Recognition (NER) models (Google-bert/bert-base-multilingual-uncased, FacebookAI/xlm-roberta-large, and Google/canine-c). A controlled experiment was then conducted to evaluate the efficacy of twelve distinct OCR-NER combinations using performance metrics such as F1 scores,Matthews Correlation Coefficient (MCC), and Structural Similarity Index (SSIM). Results demonstrated that token-based NER models successfully generalize to Swedish formats, with the combination of Tesseract and xlm-roberta-large proving the most robust, achieving the highest combined mean F1 and MCC scores before masking. However, the evaluation found that while structural image metrics like SSIM were misleadingly high (>0.9) because of naturally sparse PII, visual analysis revealed a maximum masking accuracy (F1-micro score) of approximately 85%. Ultimately, this study provides a foundational framework for localized privacy-enhancing technologies, demonstrating that open-source, multilingual models can be successfully adapted for Swedish PII without requiring extensive computational resources. However, the findings also displayed error propagation problems, where early-stage OCR tokenization errors significantly degrade downstream NER masking. This indicates that achieving fully automated redactions of highly unstructured image documents remains a highly complex challenge.
Information
- Författare
- Zeleskov, Lilia, von Trotta-Treyden, Jennifer
- Lärosäte / institution
- Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
- Publiceringsdatum
- 2026
- Uppsatstyp
- Kandidat-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Magister-uppsats, Linnéuniversitetet/Institutionen för informatik (IK)
Bekele, Wondwesen Shume, Salman Kanbar, Ahmad
Publicerad: 2026
Master-uppsats, Stockholms universitet/Institutionen för lingvistik
Meredith, Rhodri
Publicerad: 2026
Kandidat-uppsats, KTH/Hälsoinformatik och logistik
Ojanne, Benjamin, Springer, Simon
Publicerad: 2025
Kandidat-uppsats, Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Lorensson, Linus, Muhtadee, Faiyaz
Publicerad: 2026
Master-uppsats, Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
MohamadAnas, Hallak
Publicerad: 2026
Kandidat-uppsats
Björklund, Malin, Karnehed, Elvira, Lentell, Erik
Publicerad: 2026