Uppsats
Explainable AI for Automatic Document Classification in Regulated Finance
Master-uppsats
Göteborgs universitet/Institutionen för data- och informationsteknik
Publicerad: 2026-07-09
Språk: Engelska
Sammanfattning
The increasing volume of digital documents in regulated financial environmentshas created significant challenges related to information security, regulatorycompliance, and efficient information management. Financial institutions routinely process sensitive information, including internal business data, customerrecords, and regulatory documents, where incorrect handling or classificationmay result in legal, financial, and reputational consequences. Despite the importance of information classification, the process is often performed manually,making it inconsistent, time-consuming, and difficult to scale. These challengesmotivate the need for automated and trustworthy document classification systems that can support regulated organizations while maintaining transparencyand accountability.This thesis investigates the use of transformer-based language models for automatic document classification in regulated financial environments, with a particular focus on explainability and auditability. The study explores how contextualand semantic information within documents can be used to distinguish betweendifferent information sensitivity levels, including Public, Internal, Confidential,and Strictly Confidential classifications. To address privacy and regulatory constraints, the work utilizes a synthetic and semi-controlled dataset .generatedusing a controlled template-based synthetic document generation methodologywith constrained vocabulary, document structures, and contextual patterns designed to reflect the structural and linguistic characteristics of financial documents while avoiding the use of sensitive real-world data.The proposed system is based on a fine-tuned transformer architecture combined with explainable artificial intelligence (XAI) techniques. Attention-basedexplanations and Integrated Gradients feature attribution methods are integrated into the classification pipeline to provide insight into the model’s decisionmaking process. The explainability analysis investigates whether the generatedexplanations align with meaningful contextual indicators associated with document sensitivity and whether they can support transparency, trust, and compliance requirements within regulated financial settings.The experimental results demonstrate that transformer-based models can effectively learn contextual patterns related to information sensitivity within thecontrolled dataset while also providing interpretable explanations of classification decisions. The study further analyzes explanation consistency, confidencebehaviour, robustness against external documents, and potential shortcut learning effects. Since both the training and evaluation data were generated using thesame controlled template-based document generation methodology, the resultsshould be interpreted within the context of this experimental setting. Althoughseparate documents were used for training and evaluation, come from the samedataset share similar linguistic and structural characteristics. Therefore, furtherevaluation using independent datasets is required to assess the generalizability of the proposed approach.This work contributes to the growing field of explainable AI in regulated industries by demonstrating how modern natural language processing techniquescan be combined with explainability methods to support secure, transparent,and trustworthy information classification in financial organizations, while alsohighlighting the importance of independent evaluation when using controlledand synthetic data.
Information
- Författare
- Stenhammar, Zachris, Alavala, Praveen
- Lärosäte / institution
- Göteborgs universitet/Institutionen för data- och informationsteknik
- Publiceringsdatum
- 2026-07-09
- Uppsatstyp
- Master-uppsats
- Språk
- Engelska