Uppsats

Adapting Medical Language Models : A Comprehensive Approach with Domain Pretraining, Instruction Tuning, and Department-Aware Specialization

Master-uppsats

KTH/Skolan för elektroteknik och datavetenskap (EECS)

Publicerad: 2026

Språk: Engelska

Sammanfattning

Large Large Language Models (LLMs) exhibit strong general-purpose language abilities, yet adapting them to knowledge-intensive and highstakes domains such as medicine remains challenging. Medical language understanding requires not only factual knowledge but also consultation-style interaction, while clinical practice is naturally structured by departments with distinct terminology and reasoning patterns. This thesis systematically studies how to steer a pretrained foundation model toward medicine and how to further support department-level specialization in a parameter-efficient manner. We adopt a three-stage training scheme. First, we perform continued autoregressive pretraining on curated biomedical papers and medical textbooks to inject foundational medical knowledge. Second, we conduct supervised instruction fine-tuning with diverse medical supervision signals, including multiple-choice Question Answering (QA), single-turn consultations, knowledge-graph prompting, and multi-turn dialogue, to align the model with medical instructions and dialogue behavior. Third, we apply a department-aware MoELoRA adaptation stage, where department identity is used as an explicit routing signal to compose low-rank expert updates with Top-𝐾 expert selection, aiming to reduce cross-department interference while keeping the trainable parameter budget comparable to standard Low-Rank Adaptation (LoRA). We evaluate the resulting model from two complementary perspectives. On standard medical QA benchmarks, the model achieves consistent accuracy gains over the base backbone and strong open-source baselines. On dialoguecentric benchmarks, automatic metrics indicate improved reference alignment and response diversity in both single-turn and multi-turn settings. Because overlap-based metrics are insufficient to capture clinical appropriateness, we further employ rubric-based evaluation derived from Mini-CEX and its LLM-oriented extension, using Generative Pre-trained Transformer 4 (GPT- 4) as an evaluator under a structured protocol. Finally, a routing-weight case study shows clear department-dependent expert compositions, providing interpretability for the specialization mechanism. Overall, the results suggest that staged medical adaptation and departmentaware specialization can improve model performance and interpretability under benchmark-based evaluation, while further clinician-centered validation remains necessary before real-world use.

Utforska vidare

Liknande uppsatser

Uppsatser med liknande ämnen och nyckelord.