Sammanfattning

Transformer language models are poorly understood however being increasingly used in field of AI. The challenge of understanding how these models encode and process information has become critical as newer models grows in complexity and deployment. This thesis investigates if learned features within Transformer models which are stored in polysemantic superposition within Multi Layer Perceptron (MLP) layers can be disentangled into interpretable components using Sparse Autoencoders (SAEs). The thesis builds upon the foundation work of Antropic researchers and follows the methodology of their paper. The experimental framework involves training a one-layer GPT-2 like Transformer on the MiniPile dataset and collecting approximately 2 million MLP activation vectors and training sparse autoencoders, specifically encouraging sparsity and learning overcomplete representations. The results provide strong empirical evidence that most ’facts’ and semantic understanding learned by a Transformer model are stored in the MLP layer and can be unpacked using a sparse autoencoder. Qualitative analysis reveals numerous highly interpretable, monosemantic features that activate selectively on specific concepts. The findings suggest that SAEs successfully uncovers the distributed, polysemantic representations within MLP layers into more fine-grained, interpretable components. This work contributes to the growing field of mechanistic interpretability by demonstrating that sparse autoencoders approaches can reveal meaningful internal structure even in small-scale transformer models. The methodology offers a promising alternative to attention-weight analysis, directly examining what fundamental concepts models learn to represent rather than focusing on attention mechanisms. The findings have important implications for Artificial Intelligence (AI) transparency and safety, offering a concrete approach for decomposing neural network representations into more understandable components.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.