Sammanfattning

Accurate deformable registration between Computed Tomography (CT) and Magnetic Resonance Imaging (MRI) is essential for radiotherapy of mobile organs, where small misalignments at tumour–organ interfaces can propagate to planning target volume margins and affect doses delivered to organs at risk. Existing deep learning–based registration methods are fast and accurate on average, but they typically optimise surrogate similarity metrics that do not necessarily reflect how clinicians assess registration quality in critical regions. This thesis proposes and evaluates a framework that integrates reinforcement learning from human feedback (RLHF) into deep deformable CT–MRI and 4D-CT registration, so that optimisation is directly guided by expert judgement while preserving anatomical plausibility. The approach first trains a reward model that predicts an ordinal registration quality score from 3D image pairs. The reward model is trained using a combination of automatically derived proxy labels and curated ratings provided by radiation oncologists through a dedicated annotation interface. It then serves as the optimisation signal for an Advantage Actor–Critic agent, which learns to apply small residual deformations on top of a strong baseline registration. Experimental results show that the reward model provides consistent and clinically meaningful rankings of registration quality, with most errors confined to neighbouring ordinal levels. When used as a reward signal, single refinement episodes improve both the predicted quality score and intensity-based similarity measures, while maintaining Dice overlap and deformation regularity close to the baseline. Overall, these findings suggest that RLHF is a promising direction for introducing human-aligned and auditable refinements into deep deformable registration for radiotherapy.