Uppsats
Evaluating Visual-Language Models for Handwritten Text Recognition on Historical Swedish Manuscripts
Master-uppsats
Uppsala universitet/Institutionen för lingvistik och filologi
Publicerad: 2025
Språk: Engelska
Sammanfattning
Optical character recognition - and by extension, handwritten text recognition - has a long history of research and development. While text recognition for printed text has become a mature and widely adopted technology, handwritten text recognition remains a challenging task due to the high variability in human handwriting. Over time, the field has shifted from rule-based systems to more adaptable neural network-based approaches. Recently, the emergence of large language models and, subsequently, visual-language models has introduced new possibilities for tackling this problem. This thesis investigates the viability of visual-language models for handwritten text recognition on Swedish historical manuscripts, using Microsoft's Florence-2 as a case study. The model is evaluated against traditional computer vision models across individual tasks, such as text region detection and segmentation (using YOLO) and text recognition (using TrOCR). Additionally, end-to-end text recognition pipelines composed of these models are compared on text recognition accuracy. The findings show that Florence-2 performs comparably to, and in some cases surpasses, task-specific models in isolated tasks. Furthermore, a Florence-based two-step pipeline (line detection followed by text recognition) outperforms the traditional three-step pipeline (region detection, line segmentation, and text recognition) in standard text recognition metrics. However, the study also reveals that Florence-2 demands significantly greater computational resources than conventional models. As a result, while visual-language models such as Florence-2 demonstrate strong performance, at the current stage, they are not well-suited for large-scale text annotation projects. At the same time, this study is subject to several limitations, including imbalanced training data, limited exploration of multi-task training, and the use of a model not trained on a Swedish corpus. Future work may address these issues by applying data augmentation techniques to mitigate the imbalance, investigating simultaneous multi-task training strategies, and fine-tuning the language model component of the visual-language model with a corpus of historical Swedish texts.
Information
- Författare
- Pham, Hoang Ha
- Lärosäte / institution
- Uppsala universitet/Institutionen för lingvistik och filologi
- Publiceringsdatum
- 2025
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska
Utforska vidare
Liknande uppsatser
Uppsatser med liknande ämnen och nyckelord.
Master-uppsats, Mälardalens universitet/Akademin för innovation, design och teknik
Ibrahim, Mudar, Eriksson, Viktor
Publicerad: 2025
Master-uppsats, Stockholms universitet/Institutionen för data- och systemvetenskap
Suryanarayana Rao Prasanna, Navyashree
Publicerad: 2026
Master-uppsats, Lunds universitet/Matematik LTH
Erlander, Albin, Persson, Felix
Publicerad: 2024
Kandidat-uppsats, Linnéuniversitetet/Institutionen för datavetenskap och medieteknik (DM)
Cavallie Mester, Jon, Kasab, Saed
Publicerad: 2025
Kandidat-uppsats, KTH/Hälsoinformatik och logistik
Govindaraj, Gemini, Chandramohan, Moses
Publicerad: 2025
Magister-uppsats, Linnéuniversitetet/Institutionen för informatik (IK)
Bekele, Wondwesen Shume, Salman Kanbar, Ahmad
Publicerad: 2026