Uppsats

Counterfactual Explanations inLarge Language Models

Master-uppsats

Uppsala universitet/Institutionen för informationsteknologi

Publicerad: 2026

Språk: Engelska

Sammanfattning

This thesis addresses the critical need for enhanced interpretability in Large Language Models(LLMs) through counterfactual explanations. We present a modular re-implementation of IBM’s”CELL your Model” framework, designed to identify minimal input changes that induce contra-dictory LLM outputs. Our approach supports various infilling models and contradiction scoringfunctions, with experiments primarily conducted on the Phi-4 14B model. Empirical evaluations,using metrics such as flip rate and edit distance, reveal that counterfactual quality is significantlyinfluenced by infilling model performance and masking strategy. Notably, smaller LLMs exhibitlimited capability in producing coherent or contradictory outputs, and counterfactual generation issensitive to previously undocumented parameters. This work contributes a reproducible founda-tion for future research, offering critical insights into improving Explainable AI methods for LLMsand highlighting challenges in reproducibility due to inherent LLM stochasticity and undisclosedmethodological details. Because the original CELL paper was later withdrawn due to method-ological concerns, this thesis also examines reproducibility limitations and clarifies previouslyundocumented implementation details.

Information

Lärosäte / institution
Uppsala universitet/Institutionen för informationsteknologi
Publiceringsdatum
2026
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.