Uppsats

Machine Learning for Precision Medicine in Rheumatoid Arthritis

Master-uppsats

KTH/Sannolikhetsteori, matematisk fysik och statistik

Publicerad: 2025

Språk: Engelska

Sammanfattning

Rheumatoid arthritis (RA) is a chronic autoimmune disease that affects joint health and can severely impact a patient’s quality of life. Current guidelines suggest the same methotrexate (MTX) treatment plan for all newly diagnosed RA patients, despite approximately a third of patients no longer remaining on the regimen one year after treatment start. We will refer to patients who remain on the initial MTX treatment without being prescribed additional disease-modifying anti-rheumatic drugs as persistent. Previous attempts at using machine learning to predict MTX persistence at one year have yielded unsatisfactory results, with elastic net regression and random forests being two of the better models. In this project, these models were reconstructed to predict the probability of a recently diagnosed patient being MTX persistent at one year. Patients were then studied in relation to their predicted probabilities. In particular, we were interested in characterising patients consistently predicted low probabilities despite being persistent ("persistent unpredictables") and those predicted high probabilities despite not being persistent ("non-persistent un-predictables"). We also examined model calibration and attempted to characterise patients in segments of predicted probabilities that were better or worse calibrated than other segments. Attempts at charactering individuals were conducted using statistical hypothesis tests including Fisher’s exact test, the Wilcoxon rank-sum test, and the Kruskal-Wallis test to compare different subgroups. Standardised mean differences were also utilised in the first part of the analysis. Our findings showed that, compared to patients with the opposite outcome but similar predicted probabilities, "unpredictable" patients could generally not be characterised using the features available to us. However, when compared to patients with the same outcome but more accurate predictions, many features that differed significantly between the groups were identified, most notably patient age. Many of the features identified in this step were also found to be features with high influence in the model predictions, potentially explaining the differences observed. When repeating the analysis based on segments of predicted probability identified by their calibration, relatively minor changes were seen in the results. In summary, our findings point to our models lacking predictive power, poten-tially due to relevant information missing from the data or because the outcomewe modelled is not specific enough.

Information

Författare
Glimmerfors, Pia
Lärosäte / institution
KTH/Sannolikhetsteori, matematisk fysik och statistik
Publiceringsdatum
2025
Uppsatstyp
Master-uppsats
Språk
Engelska

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.