Uppsats
Debiasing LLM Annotations in Low-Sample Regimes: A Benchmarking Study
Master-uppsats
Göteborgs universitet/Institutionen för data- och informationsteknik
Publicerad: 2026-06-30
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
Recent research has shown an increasing trend toward automating data annotationusing LLMs. The main drawback of this approach is that LLM annotations sufferfrom systematic bias. Consequently, debiasing methods such as DSL and PPI++have been developed to correct for this bias. These methods have not yet beenevaluated in low-sample regimes where annotation budgets are most constrained.This thesis is a benchmarking study, evaluating these debiasing methods on fourbinary tasks, using five different LLMs for annotation, three downstream models,and varying annotation budgets between n = 20 to 200 labels.The main finding is that debiasing consistently improves model estimation errorseven at budgets as low as n = 20. The benefit is largest where the expert-only θ†baseline has the highest uncertainty. Downstream model selection is a critical factoraffecting the efficacy of debiasing methods: class prevalence and LPM yield thestrongest and most stable improvements, while logistic regression- the natural choicefor binary outcomes- introduces a persistent bias floor due to the L2 regularizationrequired in the low-sample regime.These results provide insights to help practitioners working with tight annotationbudgets, and introduce annotation informativeness as a potential pre-screeningdiagnostic for whether debiasing is likely to succeed.
Information
- Författare
- Fatima, Kesa, Törnqvist, Thea
- Lärosäte / institution
- Göteborgs universitet/Institutionen för data- och informationsteknik
- Publiceringsdatum
- 2026-06-30
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska