Sammanfattning

With the objective to classify a tabular data set of breast cancer patients with a high accuracy the self- supervised model VIME [1] is studied. The influence of several hyperparameters during pre-training is investigated and AUC of the downstream task is regarded as the measurement of performance. A larger unlabeled synthetic data set is generated using the Synthetic Data Vault (SDV) [2]. Different sizes is then pre-trained on and the result evaluated in the downstream task. Using synthetic data gives result of similar standard to the original set. Moreover an alternative mask generator implementing the correlations between features using two different methods is proposed. Both methods produce effective results compared to the original stochastic version and have arguably great potential for further research.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.