Uppsats

Automatic Anonymization of Personal Identifiable Information in Images of Unstructured Documents - In relation to Swedish-specific PII

Kandidat-uppsats

Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)

Publicerad: 2026

Språk: Engelska

Sammanfattning

In today's digital landscape, the widespread sharing of unstructured documents such as receipts and invoices presents a significant privacy risk, particularly concerning region-specific sensitive data. This bachelor's degree project investigates the automated anonymization of Swedish-specific Personal Identifiable Information (PII) within unstructured document images, addressing the knowledge gap left by existing systems trained predominantly on English-language and US-centric datasets. To tackle this, a prototype pipeline was developed using a Design Science methodology, integrating various Optical Character Recognition (OCR) tools (EasyOCR, Mindee DocTR, PaddleOCR, and Tesseract) and BERT-based/-like Named Entity Recognition (NER) models (Google-bert/bert-base-multilingual-uncased, FacebookAI/xlm-roberta-large, and Google/canine-c). A controlled experiment was then conducted to evaluate the efficacy of twelve distinct OCR-NER combinations using performance metrics such as F1 scores,Matthews Correlation Coefficient (MCC), and Structural Similarity Index (SSIM). Results demonstrated that token-based NER models successfully generalize to Swedish formats, with the combination of Tesseract and xlm-roberta-large proving the most robust, achieving the highest combined mean F1 and MCC scores before masking. However, the evaluation found that while structural image metrics like SSIM were misleadingly high (>0.9) because of naturally sparse PII, visual analysis revealed a maximum masking accuracy (F1-micro score) of approximately 85%. Ultimately, this study provides a foundational framework for localized privacy-enhancing technologies, demonstrating that open-source, multilingual models can be successfully adapted for Swedish PII without requiring extensive computational resources. However, the findings also displayed error propagation problems, where early-stage OCR tokenization errors significantly degrade downstream NER masking. This indicates that achieving fully automated redactions of highly unstructured image documents remains a highly complex challenge.

Information

Lärosäte / institution
Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Publiceringsdatum
2026
Uppsatstyp
Kandidat-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.