Uppsats
Interpreting Machine Learning Models using Conditional Counterfactual Generation
H
Chalmers tekniska högskola / Institutionen för fysik
Publicerad: 2026
Språk: Engelska
Nyckelord
klicka för att sökaSammanfattning
With the rapid development and application of complex machine learning models, theneed to interpret the internal processes of such models have become increasingly relevant.In this thesis, a novel method for interpreting black box machine learning models isproposed, where an autoencoder is used to generate reconstructions of data to visualizein an interpretable way what patterns a model has learned to detect. The method isfirst shown to work for a simple constructed problem, being able to interpret a modelthat has learned to predict the mean of an underlying normal distribution from samples.It is then evaluated for a more complex problem, where a model has learned to classifythe existence of disease in images from the CheXpert dataset of X-ray images. It isdemonstrated that naively implementing the method to interpret this model leads to theautoencoder generating adversarial patterns to trick the model, instead of showing the aninterpretable explanation of what the model has learned. To mitigate this issue, the thesisexplores adding an additional model in the latent space of the conditional autoencoder anddemonstrates that this can provide a certain degree of interpretability. Because of this,the method shows promise for interpreting black box models and with further research itmight become viable for practical use.
Information
- Författare
- Martinsson, Samuel
- Lärosäte / institution
- Chalmers tekniska högskola / Institutionen för fysik
- Publiceringsdatum
- 2026
- Uppsatstyp
- H
- Språk
- Engelska