Uppsats

Online Probability Recalibration for Machine Learning under Distribution Shift - A Study of Calibration Maintenance for In-hospital Mortality Risk Prediction

Master-uppsats

Göteborgs universitet/Institutionen för data- och informationsteknik

Publicerad: 2026-06-30

Språk: Engelska

Sammanfattning

Machine learning models used for intensive care unit (ICU) mortality predictionoutput probabilities that may guide monitoring and treatment decisions in a highstakes clinical setting. After deployment, however, distribution shift can make theseprobabilities miscalibrated, while online correction is constrained by delayed outcomelabels and a limited rolling labelled buffer.This thesis studies online post-hoc recalibration under these conditions using theeICU Collaborative Research Database. A base predictor trained on 2014 data iskept fixed, while a separate calibration layer may update during deployment on a2015 evaluation stream. Twelve update strategies are evaluated across three basemodels, four buffer sizes, and eleven shift regimes spanning the observed 2014–2015shift, which was mild in this cohort, and ten controlled score-level drift scenarios.Under the observed mild shift, no trigger-based strategy improved on fixed calibration,and S12 Adaptive Hybrid remained inactive with zero refits across all three models.Under controlled drift, fixed calibration deteriorated, while trigger-based updatingbecame beneficial. Across all eleven regimes, S12 achieved the lowest overall meanrank of 2.27, the narrowest spread of 0.92, and remained in the top three in everyregime. Under drift, it used 15–31 refits per configuration, compared with 124 forthe periodic baseline. These findings show that the value of online recalibration isregime-dependent: under mild shift, preserving the existing calibrator is sufficient,whereas under stronger or more structured drift, selective trigger-based updatingbecomes beneficial. Overall, the adaptive hybrid policy, which combines label-freeand reliability-aware signals and adapts their balance to buffer maturity, providedthe most consistent and efficient recalibration performance.These findings matter beyond this specific ICU mortality task because many deployedmachine learning systems face similar practical constraints. Predictions are availableimmediately, but reliable outcome feedback arrives late, and retraining the full modelmaybe difficult, costly, or undesirable in regulated settings. In such cases, maintainingtrustworthy probabilities is not only a matter of choosing a good calibrator, butalso of deciding when an update is justified. The results support recalibration asa selective maintenance policy, offering a practical middle ground between doingnothing and updating continuously when the future drift regime is unknown.

Information

Författare
Dakir, Halim
Lärosäte / institution
Göteborgs universitet/Institutionen för data- och informationsteknik
Publiceringsdatum
2026-06-30
Uppsatstyp
Master-uppsats
Språk
Engelska