Uppsats

Evaluating Visual-Language Models for Handwritten Text Recognition on Historical Swedish Manuscripts

Master-uppsats

Uppsala universitet/Institutionen för lingvistik och filologi

Publicerad: 2025

Språk: Engelska

Sammanfattning

Optical character recognition - and by extension, handwritten text recognition - has a long history of research and development. While text recognition for printed text has become a mature and widely adopted technology, handwritten text recognition remains a challenging task due to the high variability in human handwriting. Over time, the field has shifted from rule-based systems to more adaptable neural network-based approaches. Recently, the emergence of large language models and, subsequently, visual-language models has introduced new possibilities for tackling this problem. This thesis investigates the viability of visual-language models for handwritten text recognition on Swedish historical manuscripts, using Microsoft's Florence-2 as a case study. The model is evaluated against traditional computer vision models across individual tasks, such as text region detection and segmentation (using YOLO) and text recognition (using TrOCR). Additionally, end-to-end text recognition pipelines composed of these models are compared on text recognition accuracy. The findings show that Florence-2 performs comparably to, and in some cases surpasses, task-specific models in isolated tasks. Furthermore, a Florence-based two-step pipeline (line detection followed by text recognition) outperforms the traditional three-step pipeline (region detection, line segmentation, and text recognition) in standard text recognition metrics. However, the study also reveals that Florence-2 demands significantly greater computational resources than conventional models. As a result, while visual-language models such as Florence-2 demonstrate strong performance, at the current stage, they are not well-suited for large-scale text annotation projects. At the same time, this study is subject to several limitations, including imbalanced training data, limited exploration of multi-task training, and the use of a model not trained on a Swedish corpus. Future work may address these issues by applying data augmentation techniques to mitigate the imbalance, investigating simultaneous multi-task training strategies, and fine-tuning the language model component of the visual-language model with a corpus of historical Swedish texts.

Information

Författare
Pham, Hoang Ha
Lärosäte / institution
Uppsala universitet/Institutionen för lingvistik och filologi
Publiceringsdatum
2025
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.