Uppsats

Gaze-Driven Analysis of Expert Attention in Digital Pathology Images

Master-uppsats

Malmö universitet/Institutionen för datavetenskap och medieteknik (DVMT)

Publicerad: 2026

Språk: Engelska

Sammanfattning

Computerized KP detection in histopathological analysis plays a crucial role in evaluating Oral Squamous Cell Carcinoma. This is a common form of oral cancer, and the detection of it using deep learning methods is currently constrained by a lack of high quality annotated data. We present two artefacts. One is a methodology for collecting and pre-processing gaze-tracking data from experts during microscopical image examinations. This was done by developing SlideGazeAnnotator (SGA), a custom software for image viewing with annotation and gaze-tracking capabilities. Using SGA, we collected gaze data and evaluated pre-processing methods for it such as Velocity Threshold Identification (I-VT) and Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN). We evaluated the methodology using real gaze data that was collected by a collaborating Oral Pathologist and we found it to be a promising technique for detecting center-points of anomalies. We will be making SGA and our collected data public for future works. The second artefact is a vision-based framework that utilizes gaze data to address the problem of data scarcity using a teacher-student approach and semi-supervised learning techniques. The Teacher-Student paradigm used YOLOv3 teacher model to train pseudo-labels on an unlabeled dataset consisting of high-resolution images, thereby greatly augmenting the training set in terms of volume during the transition from pilot studies to production training manifolds. Three architectural paradigms were evaluated: YOLOv8s (Convolutional Neural Network), YOLOv10s (NMS-free Convolutional Neural Networks), and RT-DETR-L (Transformers). Results indicate that YOLOv8s remains the best performing architecture for niche medical datasets. On the other hand, transformers exhibit severe data starvation, unable to produce comparable performances despite having access to more extensive data pools. From this experiment, we conclude that even though pseudo-labeling can be used to scale datasets, CNN-based architectures with inductive bias outperform global attention-based networks with regard to datasets below the 1,000 image mark. The ultimate goal is to combine the two artefacts to lay the groundworks for a robust data generation and object detection model training pipeline for the task of anomaly detection in Pathology overall.

Information

Lärosäte / institution
Malmö universitet/Institutionen för datavetenskap och medieteknik (DVMT)
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.